{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/camex-curvature-aware-merging-of-experts","title":"CAMEx: Curvature-aware Merging of Experts","arxiv_id":"2502.18821","date":"2025-02-26","proceeding":null,"authors":["Dung V. Nguyen","Minh H. Nguyen","Luc Q. Nguyen","Rachel S. Y. Teo","Tan M. Nguyen","Linh Duy Tran"],"abstract":"Existing methods for merging experts during model training and fine-tuning predominantly rely on Euclidean geometry, which assumes a flat parameter space. This assumption can limit the model's generalization ability, especially during the pre-training phase, where the parameter manifold might exhibit more complex curvature. Curvature-aware merging methods typically require additional information and computational resources to approximate the Fisher Information Matrix, adding memory overhead. In this paper, we introduce CAMEx (\\textbf{C}urvature-\\textbf{A}ware \\textbf{M}erging of \\textbf{Ex}perts), a novel expert merging protocol that incorporates natural gradients to account for the non-Euclidean curvature of the parameter manifold. By leveraging natural gradients, CAMEx adapts more effectively to the structure of the parameter space, improving alignment between model updates and the manifold's geometry. This approach enhances both pre-training and fine-tuning, resulting in better optimization trajectories and improved generalization without the substantial memory overhead typically associated with curvature-aware methods. Our contributions are threefold: (1) CAMEx significantly outperforms traditional Euclidean-based expert merging techniques across various natural language processing tasks, leading to enhanced performance during pre-training and fine-tuning; (2) we introduce a dynamic merging architecture that optimizes resource utilization, achieving high performance while reducing computational costs, facilitating efficient scaling of large language models; and (3) we provide both theoretical and empirical evidence to demonstrate the efficiency of our proposed method.","url_abs":"https://arxiv.org/abs/2502.18821v1","url_pdf":"https://arxiv.org/pdf/2502.18821v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"camex-curvature-aware-merging-of-experts","repo_url":"https://github.com/kpup1710/CAMEx","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"jax","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2502.18821","atlas_url":"https://app.syntology.ai/?focus=2502.18821","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2502.18821"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/kpup1710/CAMEx","reach":null}],"summary":{"ran":1,"ran_draft_wrong":2},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"a1ceedbea73123a8","entry":"CAMEx","repo":"kpup1710/CAMEx","repo_kind":"official","path":"transformers/models/curve_layer_high_rank.py","file_url":"https://github.com/kpup1710/CAMEx/blob/HEAD/transformers/models/curve_layer_high_rank.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a1ceedbea73123a8"}},{"code_sha256_prefix":"9cef8b5df35244ee","entry":"_prob_in_top_k","repo":"kpup1710/CAMEx","repo_kind":"official","path":"transformers/models/curve_layer_high_rank.py","file_url":"https://github.com/kpup1710/CAMEx/blob/HEAD/transformers/models/curve_layer_high_rank.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9cef8b5df35244ee"}},{"code_sha256_prefix":"c4c0e3c114850032","entry":"noisy_top_k_gating","repo":"kpup1710/CAMEx","repo_kind":"official","path":"transformers/models/curve_layer_high_rank.py","file_url":"https://github.com/kpup1710/CAMEx/blob/HEAD/transformers/models/curve_layer_high_rank.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c4c0e3c114850032"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}