{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-31463","title":"PithTrain: A Compact and Agent-Native MoE Training System","arxiv_id":"2605.31463","date":"2026-05-29","proceeding":null,"authors":["Ruihang Lai","Hao Kang","Haozhan Tang","Akaash R. Parthasarathy","Zichun Yu","Junru Shao","Todd C. Mowry","Chenyan Xiong","Tianqi Chen"],"abstract":"Mixture-of-Experts (MoE) has become the dominant architecture for frontier language models. To meet this demand, production frameworks have built optimized MoE training stacks over years of engineering effort. Yet evolving these stacks for new architectures and system optimizations remains expensive. With the rise of AI coding agents, they could automate parts of training-framework development and accelerate this evolution. But applying them to these existing frameworks carries hidden costs, invisible to today's throughput-only evaluations. We name this missing dimension agent-task efficiency (ATE): the cost of using coding agents to understand, operate, and extend a framework. Grounded in four agent-native design principles, we build PithTrain, a compact, agent-native MoE training framework. We further introduce ATE-Bench, covering real-world training-framework tasks. Our evaluation shows PithTrain matches the throughput of production frameworks, and on ATE-Bench, PithTrain enables higher agent-task efficiency, with up to 62% fewer Agent Turns and 64% less Active GPU Time.","url_abs":"https://arxiv.org/abs/2605.31463","url_pdf":"https://arxiv.org/pdf/2605.31463","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2605.31463","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.31463"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/mlc-ai/pith-train","reach":null}],"summary":{"ran":4,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":6,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e95464673a8125b1","entry":"force_balance","repo":"mlc-ai/pith-train","repo_kind":"found_in_text","path":"pithtrain/modules/load_balance.py","file_url":"https://github.com/mlc-ai/pith-train/blob/HEAD/pithtrain/modules/load_balance.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e95464673a8125b1"}},{"code_sha256_prefix":"17790b17bcc4aa5f","entry":"muon_scale_factor","repo":"mlc-ai/pith-train","repo_kind":"found_in_text","path":"pithtrain/modules/optimizer.py","file_url":"https://github.com/mlc-ai/pith-train/blob/HEAD/pithtrain/modules/optimizer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"17790b17bcc4aa5f"}},{"code_sha256_prefix":"a627927b830319d8","entry":"strip_prefix","repo":"mlc-ai/pith-train","repo_kind":"found_in_text","path":"pithtrain/modules/checkpoint.py","file_url":"https://github.com/mlc-ai/pith-train/blob/HEAD/pithtrain/modules/checkpoint.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a627927b830319d8"}},{"code_sha256_prefix":"f1822f08ab8bef43","entry":"zeropower_via_newtonschulz5","repo":"mlc-ai/pith-train","repo_kind":"found_in_text","path":"pithtrain/modules/optimizer.py","file_url":"https://github.com/mlc-ai/pith-train/blob/HEAD/pithtrain/modules/optimizer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f1822f08ab8bef43"}},{"code_sha256_prefix":"63152bf68fc87b83","entry":"find_moe","repo":"mlc-ai/pith-train","repo_kind":"found_in_text","path":"pithtrain/modules/checkpoint.py","file_url":"https://github.com/mlc-ai/pith-train/blob/HEAD/pithtrain/modules/checkpoint.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"63152bf68fc87b83"}},{"code_sha256_prefix":"29126a25dce2a598","entry":"make_load_balance_loss_fn","repo":"mlc-ai/pith-train","repo_kind":"found_in_text","path":"pithtrain/modules/load_balance.py","file_url":"https://github.com/mlc-ai/pith-train/blob/HEAD/pithtrain/modules/load_balance.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"29126a25dce2a598"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}