Methods › Natural Language Processing › Autoencoding Transformers › MobileBERT
MobileBERT
Introduced by Zhiqing Sun et al. in MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
MobileBERT is a type of inverted-bottleneck BERT that compresses and accelerates the popular BERT model. MobileBERT is a thin version of BERT_LARGE, while equipped with bottleneck structures and a carefully designed balance between self-attentions and feed-forward networks. To train MobileBERT, we first train a specially designed teacher model, an inverted-bottleneck incorporated BERT_LARGE model. Then, we conduct knowledge transfer from this teacher to MobileBERT. Like the original BERT, MobileBERT is task-agnostic, that is, it can be generically applied to various downstream NLP tasks via simple fine-tuning. It is trained by layer-to-layer imitating the inverted bottleneck BERT.
Papers archive 2025-07-28
12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Efficient Intent-Based Filtering for Multi-Party Conversations Using Knowledge Distillation from LLMs 21 Mar 2025 · 0 repositories · arXiv:2503.17336
-
FedMentalCare: Towards Privacy-Preserving Fine-Tuned LLMs to Analyze Mental Health Status Using Federated Learning Framework 27 Feb 2025 · 0 repositories · arXiv:2503.05786
-
Resource-Efficient Transformer Architecture: Optimizing Memory and Execution Time for Real-Time Applications 25 Dec 2024 · 0 repositories · arXiv:2501.00042
-
Efficient Deployment of Transformer Models in Analog In-Memory Computing Hardware 26 Nov 2024 · 1 repository · arXiv:2411.17367
-
On-Device Emoji Classifier Trained with GPT-based Data Augmentation for a Mobile Keyboard 6 Nov 2024 · 0 repositories · arXiv:2411.05031
-
PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT 11 Aug 2024 · 2 repositories · arXiv:2408.05667
-
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow 5 Aug 2024 · 1 repository · arXiv:2408.02473
-
Quantized Transformer Language Model Implementations on Edge Devices 6 Oct 2023 · 0 repositories · arXiv:2310.03971
-
AutoDistill: an End-to-End Framework to Explore and Distill Hardware-Efficient Language Models 21 Jan 2022 · 0 repositories · arXiv:2201.08539
-
Character-level HyperNetworks for Hate Speech Detection 11 Nov 2021 · 1 repository · arXiv:2111.06336
-
LIDSNet: A Lightweight on-device Intent Detection model using Deep Siamese Network 6 Oct 2021 · 0 repositories · arXiv:2110.15717
-
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices 6 Apr 2020 · 7 repositories · arXiv:2004.02984
Tasks archive 2025-07-28
20 shown of 24 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections