{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-does-critical-batch-size-scale-in-pre","title":"How Does Critical Batch Size Scale in Pre-training?","arxiv_id":"2410.21676","date":"2024-10-29","proceeding":null,"authors":["HANLIN ZHANG","Depen Morwani","Nikhil Vyas","Jingfeng Wu","Difan Zou","Udaya Ghai","Dean Foster","Sham Kakade"],"abstract":"Training large-scale models under given resources requires careful design of parallelism strategies. In particular, the efficiency notion of critical batch size (CBS), concerning the compromise between time and compute, marks the threshold beyond which greater data parallelism leads to diminishing returns. To operationalize it, we propose a measure of CBS and pre-train a series of auto-regressive language models, ranging from 85 million to 1.2 billion parameters, on the C4 dataset. Through extensive hyper-parameter sweeps and careful control of factors such as batch size, momentum, and learning rate along with its scheduling, we systematically investigate the impact of scale on CBS. Then we fit scaling laws with respect to model and data sizes to decouple their effects. Overall, our results demonstrate that CBS scales primarily with data size rather than model size, a finding we justify theoretically through the analysis of infinite-width limits of neural networks and infinite-dimensional least squares regression. Of independent interest, we highlight the importance of common hyper-parameter choices and strategies for studying large-scale pre-training beyond fixed training durations.","url_abs":"https://arxiv.org/abs/2410.21676v4","url_pdf":"https://arxiv.org/pdf/2410.21676v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-does-critical-batch-size-scale-in-pre","repo_url":"https://github.com/hlzhang109/critical-batch-size","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"scheduling","task_name":"Scheduling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.21676","atlas_url":"https://app.syntology.ai/?focus=2410.21676","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.21676"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hlzhang109/critical-batch-size","reach":null}],"summary":{"ran":2,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d15280161ba29860","entry":"SpeedMonitor","repo":"hlzhang109/critical-batch-size","repo_kind":"official","path":"olmo/trainer.py","file_url":"https://github.com/hlzhang109/critical-batch-size/blob/HEAD/olmo/trainer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d15280161ba29860"}},{"code_sha256_prefix":"de29816c9d3d62df","entry":"SpeedMonitorConfig","repo":"hlzhang109/critical-batch-size","repo_kind":"official","path":"olmo/trainer.py","file_url":"https://github.com/hlzhang109/critical-batch-size/blob/HEAD/olmo/trainer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"de29816c9d3d62df"}},{"code_sha256_prefix":"f840835293d978b8","entry":"BaseConfig","repo":"hlzhang109/critical-batch-size","repo_kind":"official","path":"olmo/trainer.py","file_url":"https://github.com/hlzhang109/critical-batch-size/blob/HEAD/olmo/trainer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f840835293d978b8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}