{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2601-15021","title":"Mixture-of-Experts Models in Vision: Routing, Optimization, and Generalization","arxiv_id":"2601.15021","date":"2026-01-21","proceeding":null,"authors":["Adam Rokah","Daniel Veress","Caleb Caulk","Sourav Sharan"],"abstract":"Mixture-of-Experts (MoE) architectures enable conditional computation by routing inputs to multiple expert subnetworks and are often motivated as a mechanism for scaling large language models. In this project, we instead study MoE behavior in an image classification setting, focusing on predictive performance, expert utilization, and generalization. We compare dense, SoftMoE, and SparseMoE classifier heads on the CIFAR10 dataset under comparable model capacity. Both MoE variants achieve slightly higher validation accuracy than the dense baseline while maintaining balanced expert utilization through regularization, avoiding expert collapse. To analyze generalization, we compute Hessian-based sharpness metrics at convergence, including the largest eigenvalue and trace of the loss Hessian, evaluated on both training and test data. We find that SoftMoE exhibits higher sharpness by these metrics, while Dense and SparseMoE lie in a similar curvature regime, despite all models achieving comparable generalization performance. Complementary loss surface perturbation analyses reveal qualitative differences in non-local behavior under finite parameter perturbations between dense and MoE models, which help contextualize curvature-based measurements without directly explaining validation accuracy. We further evaluate empirical inference efficiency and show that naively implemented conditional routing does not yield inference speedups on modern hardware at this scale, highlighting the gap between theoretical and realized efficiency in sparse MoE models.","url_abs":"https://arxiv.org/abs/2601.15021","url_pdf":"https://arxiv.org/pdf/2601.15021","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2601.15021","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2601.15021"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/moe-project-uu/mixture-of-experts-project","reach":null}],"summary":{"unverified":5},"by_repo_kind":{"found_in_text":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"474043b1e52a6fb9","entry":"build_head","repo":"moe-project-uu/mixture-of-experts-project","repo_kind":"found_in_text","path":"src/moe/heads/factory.py","file_url":"https://github.com/moe-project-uu/mixture-of-experts-project/blob/HEAD/src/moe/heads/factory.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"474043b1e52a6fb9"}},{"code_sha256_prefix":"63b556bf3389cb20","entry":"make_subset_loader","repo":"moe-project-uu/mixture-of-experts-project","repo_kind":"found_in_text","path":"src/moe/utils/helpers.py","file_url":"https://github.com/moe-project-uu/mixture-of-experts-project/blob/HEAD/src/moe/utils/helpers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"63b556bf3389cb20"}},{"code_sha256_prefix":"32b7ffaaeb746350","entry":"shazeer_importance_loss","repo":"moe-project-uu/mixture-of-experts-project","repo_kind":"found_in_text","path":"src/moe/utils/losses.py","file_url":"https://github.com/moe-project-uu/mixture-of-experts-project/blob/HEAD/src/moe/utils/losses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"32b7ffaaeb746350"}},{"code_sha256_prefix":"52965e5e24e2d339","entry":"shazeer_load_loss","repo":"moe-project-uu/mixture-of-experts-project","repo_kind":"found_in_text","path":"src/moe/utils/losses.py","file_url":"https://github.com/moe-project-uu/mixture-of-experts-project/blob/HEAD/src/moe/utils/losses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"52965e5e24e2d339"}},{"code_sha256_prefix":"8e92a34d56fa9769","entry":"softmoe_load_balance","repo":"moe-project-uu/mixture-of-experts-project","repo_kind":"found_in_text","path":"src/moe/utils/losses.py","file_url":"https://github.com/moe-project-uu/mixture-of-experts-project/blob/HEAD/src/moe/utils/losses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8e92a34d56fa9769"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}