{"url":"/method/switch-ffn","slug":"switch-ffn","name":"Switch FFN","full_name":"Switch FFN","full_name_withheld":false,"description_markdown":"A **Switch FFN** is a sparse layer that operates independently on tokens within an input sequence. It is shown in the blue block in the figure. We diagram two tokens ($x\\_{1}$ = “More” and $x\\_{2}$ = “Parameters” below) being routed (solid lines) across four FFN experts, where the router independently routes each token. The switch FFN layer returns the output of the selected FFN multiplied by the router gate value (dotted-line).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","paper":"/paper/switch-transformers-scaling-to-trillion","first_author":"William Fedus","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/switch-transformers-scaling-to-trillion"},"source":{"url":"https://arxiv.org/abs/2101.03961v3","title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Feedforward Networks","url":"/methods/category/feedforward-networks","pwc_aliases":[]}],"n_papers_tagged":20,"archive_num_papers":20,"papers_newest_first":[{"paper":null,"title":"QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration","date":"2025-05-10","arxiv_id":"2505.06481","n_code_links":0,"syntology":null},{"paper":null,"title":"ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses","date":"2025-03-23","arxiv_id":"2504.08744","n_code_links":0,"syntology":null},{"paper":"/paper/resmoe-space-efficient-compression-of-mixture","title":"ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration","date":"2025-03-10","arxiv_id":"2503.06881","n_code_links":1,"syntology":{"ran":5,"of":15,"unverified":10,"pointer_only":0}},{"paper":null,"title":"Sparse Backpropagation for MoE Training","date":"2023-10-01","arxiv_id":"2310.00811","n_code_links":0,"syntology":null},{"paper":null,"title":"SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget","date":"2023-08-29","arxiv_id":"2308.15030","n_code_links":0,"syntology":null},{"paper":"/paper/condensing-multilingual-knowledge-with","title":"Condensing Multilingual Knowledge with Lightweight Language-Specific Modules","date":"2023-05-23","arxiv_id":"2305.13993","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language Model","date":"2023-05-23","arxiv_id":"2305.13999","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset","date":"2023-03-22","arxiv_id":"2303.12892","n_code_links":0,"syntology":null},{"paper":null,"title":"SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing","date":"2022-12-10","arxiv_id":"2212.05191","n_code_links":0,"syntology":null},{"paper":"/paper/argumentative-text-generation-in-economic","title":"Argumentative Text Generation in Economic Domain","date":"2022-06-18","arxiv_id":"2206.09251","n_code_links":1,"syntology":null},{"paper":null,"title":"Automatic Summarization of Russian Texts: Comparison of Extractive and Abstractive Methods","date":"2022-06-18","arxiv_id":"2206.09253","n_code_links":0,"syntology":null},{"paper":"/paper/build-a-robust-qa-system-with-transformer","title":"Build a Robust QA System with Transformer-based Mixture of Experts","date":"2022-03-20","arxiv_id":"2204.09598","n_code_links":1,"syntology":null},{"paper":"/paper/efficient-language-modeling-with-sparse-all","title":"Efficient Language Modeling with Sparse all-MLP","date":"2022-03-14","arxiv_id":"2203.06850","n_code_links":0,"syntology":null},{"paper":null,"title":"Switch Trajectory Transformer with Distributional Value Approximation for Multi-Task Reinforcement Learning","date":"2022-03-14","arxiv_id":"2203.07413","n_code_links":0,"syntology":null},{"paper":null,"title":"Mixture-of-Experts with Expert Choice Routing","date":"2022-02-18","arxiv_id":"2202.09368","n_code_links":0,"syntology":null},{"paper":null,"title":"M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining","date":"2021-10-08","arxiv_id":"2110.03888","n_code_links":0,"syntology":null},{"paper":"/paper/taming-sparsely-activated-transformer-with","title":"Taming Sparsely Activated Transformer with Stochastic Experts","date":"2021-10-08","arxiv_id":"2110.04260","n_code_links":1,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":null,"title":"Random Offset Block Embedding Array (ROBE) for CriteoTB Benchmark MLPerf DLRM Model : 1000$\\times$ Compression and 3.1$\\times$ Faster Inference","date":"2021-08-04","arxiv_id":"2108.02191","n_code_links":0,"syntology":null},{"paper":null,"title":"Carbon Emissions and Large Neural Network Training","date":"2021-04-21","arxiv_id":"2104.10350","n_code_links":0,"syntology":null},{"paper":"/paper/switch-transformers-scaling-to-trillion","title":"Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity","date":"2021-01-11","arxiv_id":"2101.03961","n_code_links":8,"syntology":{"ran":7,"of":13,"unverified":6,"pointer_only":8}}],"papers_shown":20,"tasks":[{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":13},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":null,"name":"GPU","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/machine-translation","name":"Machine Translation","papers":3},{"task":"/task/question-answering","name":"Question Answering","papers":3},{"task":"/task/text-generation","name":"Text Generation","papers":2},{"task":"/task/translation","name":"Translation","papers":2},{"task":"/task/all","name":"All","papers":1},{"task":null,"name":"Avg","papers":1},{"task":null,"name":"CPU","papers":1},{"task":"/task/common-sense-reasoning","name":"Common Sense Reasoning","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":1},{"task":"/task/large-language-model","name":"Large Language Model","papers":1},{"task":"/task/model-compression","name":"Model Compression","papers":1},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/rag","name":"RAG","papers":1},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":1}],"tasks_shown":20,"n_tasks":30,"usage_by_year":[{"year":"2021","papers":5},{"year":"2022","papers":7},{"year":"2023","papers":5},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/switch-ffn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}