Methods › Natural Language Processing › Transformers › Switch Transformer
Switch Transformer
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Switch Transformer is a sparsely-activated expert Transformer model that aims to simplify and improve over Mixture of Experts. Through distillation of sparse pre-trained and specialized fine-tuned models into small dense models, it reduces the model size by up to 99% while preserving 30% of the quality gains of the large sparse teacher. It also uses selective precision training that enables training with lower bfloat16 precision, as well as an initialization scheme that allows for scaling to a larger number of experts, and also increased regularization that improves sparse model fine-tuning and multi-task training.
Papers archive 2025-07-28
20 shown of 20, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration 10 May 2025 · 0 repositories · arXiv:2505.06481
-
ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses 23 Mar 2025 · 0 repositories · arXiv:2504.08744
-
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration 10 Mar 2025 · 1 repository · arXiv:2503.06881Syntology ran 5 of 15 samples · 10 unverified
-
Sparse Backpropagation for MoE Training 1 Oct 2023 · 0 repositories · arXiv:2310.00811
-
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget 29 Aug 2023 · 0 repositories · arXiv:2308.15030
-
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules 23 May 2023 · 1 repository · arXiv:2305.13993Syntology ran 2 of 2 samples · 0 unverified
-
Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language Model 23 May 2023 · 0 repositories · arXiv:2305.13999
-
Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset 22 Mar 2023 · 0 repositories · arXiv:2303.12892
-
SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing 10 Dec 2022 · 0 repositories · arXiv:2212.05191
-
Argumentative Text Generation in Economic Domain 18 Jun 2022 · 1 repository · arXiv:2206.09251
-
Automatic Summarization of Russian Texts: Comparison of Extractive and Abstractive Methods 18 Jun 2022 · 0 repositories · arXiv:2206.09253
-
Build a Robust QA System with Transformer-based Mixture of Experts 20 Mar 2022 · 1 repository · arXiv:2204.09598
-
Efficient Language Modeling with Sparse all-MLP 14 Mar 2022 · 0 repositories · arXiv:2203.06850
-
Switch Trajectory Transformer with Distributional Value Approximation for Multi-Task Reinforcement Learning 14 Mar 2022 · 0 repositories · arXiv:2203.07413
-
Mixture-of-Experts with Expert Choice Routing 18 Feb 2022 · 0 repositories · arXiv:2202.09368
-
M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining 8 Oct 2021 · 0 repositories · arXiv:2110.03888
-
Taming Sparsely Activated Transformer with Stochastic Experts 8 Oct 2021 · 1 repository · arXiv:2110.04260Syntology ran 0 of 2 samples · 2 unverified
-
Random Offset Block Embedding Array (ROBE) for CriteoTB Benchmark MLPerf DLRM Model : 1000× Compression and 3.1× Faster Inference 4 Aug 2021 · 0 repositories · arXiv:2108.02191
-
Carbon Emissions and Large Neural Network Training 21 Apr 2021 · 0 repositories · arXiv:2104.10350
-
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity 11 Jan 2021 · 8 repositories · arXiv:2101.03961Syntology ran 7 of 13 samples · 6 unverified · 8 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 30 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections