{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pushing-mixture-of-experts-to-the-limit","title":"Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning","arxiv_id":"2309.05444","date":"2023-09-11","proceeding":null,"authors":["Ted Zadouri","Ahmet Üstün","Arash Ahmadian","Beyza Ermiş","Acyr Locatelli","Sara Hooker"],"abstract":"The Mixture of Experts (MoE) is a widely known neural architecture where an ensemble of specialized sub-models optimizes overall performance with a constant computational cost. However, conventional MoEs pose challenges at scale due to the need to store all experts in memory. In this paper, we push MoE to the limit. We propose extremely parameter-efficient MoE by uniquely combining MoE architecture with lightweight experts.Our MoE architecture outperforms standard parameter-efficient fine-tuning (PEFT) methods and is on par with full fine-tuning by only updating the lightweight experts -- less than 1% of an 11B parameters model. Furthermore, our method generalizes to unseen tasks as it does not depend on any prior task knowledge. Our research underscores the versatility of the mixture of experts architecture, showcasing its ability to deliver robust performance even when subjected to rigorous parameter constraints. Our code used in all the experiments is publicly available here: https://github.com/for-ai/parameter-efficient-moe.","url_abs":"https://arxiv.org/abs/2309.05444v1","url_pdf":"https://arxiv.org/pdf/2309.05444v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pushing-mixture-of-experts-to-the-limit","repo_url":"https://github.com/for-ai/parameter-efficient-moe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok"}}],"tasks":[{"task_slug":"mixture-of-experts","task_name":"Mixture-of-Experts"},{"task_slug":"parameter-efficient-fine-tuning","task_name":"parameter-efficient fine-tuning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2309.05444","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.05444"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/for-ai/parameter-efficient-moe","reach":{"status":"ok"}}],"summary":{"ran":4},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"b7af70f17bfb5e07","entry":"feature_to_spec","repo":"for-ai/parameter-efficient-moe","repo_kind":"official","path":"t0_data/utils.py","file_url":"https://github.com/for-ai/parameter-efficient-moe/blob/HEAD/t0_data/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b7af70f17bfb5e07"}},{"code_sha256_prefix":"de62b3eab371bdf6","entry":"hf_dataset_to_tf_dataset","repo":"for-ai/parameter-efficient-moe","repo_kind":"official","path":"t0_data/utils.py","file_url":"https://github.com/for-ai/parameter-efficient-moe/blob/HEAD/t0_data/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"de62b3eab371bdf6"}},{"code_sha256_prefix":"c22561ab35c12af3","entry":"match_any","repo":"for-ai/parameter-efficient-moe","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/for-ai/parameter-efficient-moe/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c22561ab35c12af3"}},{"code_sha256_prefix":"11b64b0a85dfb964","entry":"strip_whitespace","repo":"for-ai/parameter-efficient-moe","repo_kind":"official","path":"t0_data/tasks.py","file_url":"https://github.com/for-ai/parameter-efficient-moe/blob/HEAD/t0_data/tasks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"11b64b0a85dfb964"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}