{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/yuan-2-0-m32-mixture-of-experts-with","title":"Yuan 2.0-M32: Mixture of Experts with Attention Router","arxiv_id":"2405.17976","date":"2024-05-28","proceeding":null,"authors":["Shaohua Wu","Jiangang Luo","Xi Chen","Lingjun Li","Xudong Zhao","Tong Yu","Chao Wang","Yue Wang","Fei Wang","Weixu Qiao","Houbo He","Zeru Zhang","Zeyu Sun","Junxiong Mao","Chong Shen"],"abstract":"Yuan 2.0-M32, with a similar base architecture as Yuan-2.0 2B, uses a mixture-of-experts architecture with 32 experts of which 2 experts are active. A new router network, Attention Router, is proposed and adopted for a more efficient selection of experts, which improves the accuracy compared to the model with classical router network. Yuan 2.0-M32 is trained with 2000B tokens from scratch, and the training computation consumption is only 9.25% of a dense model at the same parameter scale. Yuan 2.0-M32 demonstrates competitive capability on coding, math, and various domains of expertise, with only 3.7B active parameters of 40B in total, and 7.4 GFlops forward computation per token, both of which are only 1/19 of Llama3-70B. Yuan 2.0-M32 surpass Llama3-70B on MATH and ARC-Challenge benchmark, with accuracy of 55.89 and 95.8 respectively. The models and source codes of Yuan 2.0-M32 are released at Github1.","url_abs":"https://arxiv.org/abs/2405.17976v2","url_pdf":"https://arxiv.org/pdf/2405.17976v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"yuan-2-0-m32-mixture-of-experts-with","repo_url":"https://github.com/ieit-yuan/yuan2.0-m32","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"arc","task_name":"ARC"},{"task_slug":"math","task_name":"Math"},{"task_slug":"mixture-of-experts","task_name":"Mixture-of-Experts"}],"methods":[{"method_slug":"base","method_name":"BASE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2405.17976","atlas_url":"https://app.syntology.ai/?focus=2405.17976","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2405.17976"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ieit-yuan/yuan2.0-m32","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1f4447581128a711","entry":"get_batch","repo":"ieit-yuan/yuan2.0-m32","repo_kind":"official","path":"pretrain_vision_classify.py","file_url":"https://github.com/ieit-yuan/yuan2.0-m32/blob/HEAD/pretrain_vision_classify.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"1f4447581128a711"}},{"code_sha256_prefix":"92287bdc91b8ac90","entry":"get_batch","repo":"ieit-yuan/yuan2.0-m32","repo_kind":"official","path":"pretrain_vision_dino.py","file_url":"https://github.com/ieit-yuan/yuan2.0-m32/blob/HEAD/pretrain_vision_dino.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"92287bdc91b8ac90"}},{"code_sha256_prefix":"edf213e729a014e0","entry":"get_batch","repo":"ieit-yuan/yuan2.0-m32","repo_kind":"official","path":"pretrain_vision_inpaint.py","file_url":"https://github.com/ieit-yuan/yuan2.0-m32/blob/HEAD/pretrain_vision_inpaint.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"edf213e729a014e0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}