{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/not-just-scaling-laws-towards-a-better","title":"Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions","arxiv_id":"2503.03862","date":"2025-03-05","proceeding":null,"authors":["Emmy Liu","Amanda Bertsch","Lintang Sutawika","Lindia Tjuatja","Patrick Fernandes","Lara Marinov","Michael Chen","Shreya Singhal","Carolin Lawrence","aditi raghunathan","Kiril Gashteovski","Graham Neubig"],"abstract":"Improvements in language model capabilities are often attributed to increasing model size or training data, but in some cases smaller models trained on curated data or with different architectural decisions can outperform larger ones trained on more tokens. What accounts for this? To quantify the impact of these design choices, we meta-analyze 92 open-source pretrained models across a wide array of scales, including state-of-the-art open-weights models as well as less performant models and those with less conventional design decisions. We find that by incorporating features besides model size and number of training tokens, we can achieve a relative 3-28% increase in ability to predict downstream performance compared with using scale alone. Analysis of model design decisions reveal insights into data composition, such as the trade-off between language and code tasks at 15-25\\% code, as well as the better performance of some architectural decisions such as choosing rotary over learned embeddings. Broadly, our framework lays a foundation for more systematic investigation of how model development choices shape final capabilities.","url_abs":"https://arxiv.org/abs/2503.03862v1","url_pdf":"https://arxiv.org/pdf/2503.03862v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"not-just-scaling-laws-towards-a-better","repo_url":"https://github.com/nightingal3/llm-pretraining-behaviours","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.03862","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.03862"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nightingal3/llm-pretraining-behaviours","reach":null}],"summary":{"ran_draft_wrong":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"489103a55ac3180f","entry":"prepare_task_data","repo":"nightingal3/llm-pretraining-behaviours","repo_kind":"official","path":"performance_prediction/performance_predict_from_db.py","file_url":"https://github.com/nightingal3/llm-pretraining-behaviours/blob/HEAD/performance_prediction/performance_predict_from_db.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"489103a55ac3180f"}},{"code_sha256_prefix":"c2e9a027226aca3d","entry":"preprocess_data","repo":"nightingal3/llm-pretraining-behaviours","repo_kind":"official","path":"performance_prediction/performance_predict_from_db.py","file_url":"https://github.com/nightingal3/llm-pretraining-behaviours/blob/HEAD/performance_prediction/performance_predict_from_db.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c2e9a027226aca3d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}