Methods › Natural Language Processing › Autoencoding Transformers › MobileBERT

MobileBERT

12 papers tagged archive 2025-07-28

Introduced by Zhiqing Sun et al. in MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

MobileBERT is a type of inverted-bottleneck BERT that compresses and accelerates the popular BERT model. MobileBERT is a thin version of BERT_LARGE, while equipped with bottleneck structures and a carefully designed balance between self-attentions and feed-forward networks. To train MobileBERT, we first train a specially designed teacher model, an inverted-bottleneck incorporated BERT_LARGE model. Then, we conduct knowledge transfer from this teacher to MobileBERT. Like the original BERT, MobileBERT is task-agnostic, that is, it can be generically applied to various downstream NLP tasks via simple fine-tuning. It is trained by layer-to-layer imitating the inverted bottleneck BERT.

PaperSource

Papers archive 2025-07-28

12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 24 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Computational Efficiency2
Data Augmentation2
Knowledge Distillation2
Language Modelling2
Privacy Preserving2
Transfer Learning2
Bayesian Optimization1
Federated Learning1
Hate Speech Detection1
Intent Classification1
Intent Detection1
Language Modeling1
Large Language Model1
Model Compression1
Natural Language Inference1
Natural Language Understanding1
Neural Architecture Search1
Phishing Website Detection1
Quantization1
Question Answering1

Usage over time archive 2025-07-28

Papers per year tagged with MobileBERT: 2020 to 2025, peak 5 5 0 2020: 1 paper 2020 2021: 2 papers 2021 2022: 1 paper 2022 2023: 1 paper 2023 2024: 5 papers 2024 2025: 2 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (12 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Autoencoding TransformersTransformers

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections