{"url":"/method/switch-transformer","slug":"switch-transformer","name":"Switch Transformer","full_name":"Switch Transformer","full_name_withheld":false,"description_markdown":"**Switch Transformer** is a sparsely-activated expert [Transformer](https://paperswithcode.com/methods/category/transformers) model that aims to simplify and improve over Mixture of Experts. Through distillation of sparse pre-trained and specialized fine-tuned models into small dense models, it reduces the model size by up to 99% while preserving 30% of the quality gains of the large sparse teacher. It also uses selective precision training that enables training with lower bfloat16 precision, as well as an initialization scheme that allows for scaling to a larger number of experts, and also increased regularization that improves sparse model fine-tuning and multi-task training.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2101.03961v3","title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":20,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration","date":"2025-05-10","arxiv_id":"2505.06481","n_code_links":0,"syntology":null},{"paper":null,"title":"ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses","date":"2025-03-23","arxiv_id":"2504.08744","n_code_links":0,"syntology":null},{"paper":"/paper/resmoe-space-efficient-compression-of-mixture","title":"ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration","date":"2025-03-10","arxiv_id":"2503.06881","n_code_links":1,"syntology":{"ran":5,"of":15,"unverified":10,"pointer_only":0}},{"paper":null,"title":"Sparse Backpropagation for MoE Training","date":"2023-10-01","arxiv_id":"2310.00811","n_code_links":0,"syntology":null},{"paper":null,"title":"SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget","date":"2023-08-29","arxiv_id":"2308.15030","n_code_links":0,"syntology":null},{"paper":"/paper/condensing-multilingual-knowledge-with","title":"Condensing Multilingual Knowledge with Lightweight Language-Specific Modules","date":"2023-05-23","arxiv_id":"2305.13993","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language Model","date":"2023-05-23","arxiv_id":"2305.13999","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset","date":"2023-03-22","arxiv_id":"2303.12892","n_code_links":0,"syntology":null},{"paper":null,"title":"SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing","date":"2022-12-10","arxiv_id":"2212.05191","n_code_links":0,"syntology":null},{"paper":"/paper/argumentative-text-generation-in-economic","title":"Argumentative Text Generation in Economic Domain","date":"2022-06-18","arxiv_id":"2206.09251","n_code_links":1,"syntology":null},{"paper":null,"title":"Automatic Summarization of Russian Texts: Comparison of Extractive and Abstractive Methods","date":"2022-06-18","arxiv_id":"2206.09253","n_code_links":0,"syntology":null},{"paper":"/paper/build-a-robust-qa-system-with-transformer","title":"Build a Robust QA System with Transformer-based Mixture of Experts","date":"2022-03-20","arxiv_id":"2204.09598","n_code_links":1,"syntology":null},{"paper":"/paper/efficient-language-modeling-with-sparse-all","title":"Efficient Language Modeling with Sparse all-MLP","date":"2022-03-14","arxiv_id":"2203.06850","n_code_links":0,"syntology":null},{"paper":null,"title":"Switch Trajectory Transformer with Distributional Value Approximation for Multi-Task Reinforcement Learning","date":"2022-03-14","arxiv_id":"2203.07413","n_code_links":0,"syntology":null},{"paper":null,"title":"Mixture-of-Experts with Expert Choice Routing","date":"2022-02-18","arxiv_id":"2202.09368","n_code_links":0,"syntology":null},{"paper":null,"title":"M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining","date":"2021-10-08","arxiv_id":"2110.03888","n_code_links":0,"syntology":null},{"paper":"/paper/taming-sparsely-activated-transformer-with","title":"Taming Sparsely Activated Transformer with Stochastic Experts","date":"2021-10-08","arxiv_id":"2110.04260","n_code_links":1,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":null,"title":"Random Offset Block Embedding Array (ROBE) for CriteoTB Benchmark MLPerf DLRM Model : 1000$\\times$ Compression and 3.1$\\times$ Faster Inference","date":"2021-08-04","arxiv_id":"2108.02191","n_code_links":0,"syntology":null},{"paper":null,"title":"Carbon Emissions and Large Neural Network Training","date":"2021-04-21","arxiv_id":"2104.10350","n_code_links":0,"syntology":null},{"paper":"/paper/switch-transformers-scaling-to-trillion","title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","date":"2021-01-11","arxiv_id":"2101.03961","n_code_links":8,"syntology":{"ran":7,"of":13,"unverified":6,"pointer_only":8}}],"papers_shown":20,"tasks":[{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":13},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":null,"name":"GPU","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/machine-translation","name":"Machine Translation","papers":3},{"task":"/task/question-answering","name":"Question Answering","papers":3},{"task":"/task/text-generation","name":"Text Generation","papers":2},{"task":"/task/translation","name":"Translation","papers":2},{"task":"/task/all","name":"All","papers":1},{"task":null,"name":"Avg","papers":1},{"task":null,"name":"CPU","papers":1},{"task":"/task/common-sense-reasoning","name":"Common Sense Reasoning","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":1},{"task":"/task/large-language-model","name":"Large Language Model","papers":1},{"task":"/task/model-compression","name":"Model Compression","papers":1},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/rag","name":"RAG","papers":1},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":1}],"tasks_shown":20,"n_tasks":30,"usage_by_year":[{"year":"2021","papers":5},{"year":"2022","papers":7},{"year":"2023","papers":5},{"year":"2025","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/switch-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}