{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-empirical-model-of-large-batch-training","title":"An Empirical Model of Large-Batch Training","arxiv_id":"1812.06162","date":"2018-12-14","proceeding":null,"authors":["Sam McCandlish","Jared Kaplan","Dario Amodei","OpenAI Dota Team"],"abstract":"In an increasing number of domains it has been demonstrated that deep\nlearning models can be trained using relatively large batch sizes without\nsacrificing data efficiency. However the limits of this massive data\nparallelism seem to differ from domain to domain, ranging from batches of tens\nof thousands in ImageNet to batches of millions in RL agents that play the game\nDota 2. To our knowledge there is limited conceptual understanding of why these\nlimits to batch size differ or how we might choose the correct batch size in a\nnew domain. In this paper, we demonstrate that a simple and easy-to-measure\nstatistic called the gradient noise scale predicts the largest useful batch\nsize across many domains and applications, including a number of supervised\nlearning datasets (MNIST, SVHN, CIFAR-10, ImageNet, Billion Word),\nreinforcement learning domains (Atari and Dota), and even generative model\ntraining (autoencoders on SVHN). We find that the noise scale increases as the\nloss decreases over a training run and depends on the model size primarily\nthrough improved model performance. Our empirically-motivated theory also\ndescribes the tradeoff between compute-efficiency and time-efficiency, and\nprovides a rough model of the benefits of adaptive batch-size training.","url_abs":"http://arxiv.org/abs/1812.06162v1","url_pdf":"http://arxiv.org/pdf/1812.06162v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/RexGLiu/rlpyt_crbp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/akterskii/rlpyt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/alexsax/testing_rlpyt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"gone","observed_at":"2026-09-17","how":"tree_404+repo_404"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/astooke/rlpyt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/eac-replication/eac-replication","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/hal-314/fastai-batch-size-finder","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/petuum/adaptdl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/sandeeprockstar/Pose_Estimation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/sarahisyoung/rlpyt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-empirical-model-of-large-batch-training","repo_url":"https://github.com/DanyWind/fastai_bs_finder","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"dota-2","task_name":"Dota 2"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1812.06162","atlas_url":"https://app.syntology.ai/?focus=1812.06162","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1812.06162"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/akterskii/rlpyt","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hal-314/fastai-batch-size-finder","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/astooke/rlpyt","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/petuum/adaptdl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sarahisyoung/rlpyt","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eac-replication/eac-replication","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alexsax/testing_rlpyt","reach":{"status":"gone","observed_at":"2026-09-17","how":"tree_404+repo_404"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/RexGLiu/rlpyt_crbp","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/DanyWind/fastai_bs_finder","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sandeeprockstar/Pose_Estimation","reach":{"status":"ok"}}],"summary":{"unverified":6},"by_repo_kind":{"listed":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a862dac46152a1e9","entry":"compute_big_batch_gradient","repo":"davidandym/task-conflict-in-text-to-text-learners","repo_kind":"listed","path":"src/measure_conflict.py","file_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners/blob/HEAD/src/measure_conflict.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a862dac46152a1e9"}},{"code_sha256_prefix":"357c3ce4115b3f1b","entry":"get_conflict_metric_keys","repo":"davidandym/task-conflict-in-text-to-text-learners","repo_kind":"listed","path":"src/measure_conflict.py","file_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners/blob/HEAD/src/measure_conflict.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"357c3ce4115b3f1b"}},{"code_sha256_prefix":"854be3983683cc44","entry":"masked_softmax","repo":"davidandym/task-conflict-in-text-to-text-learners","repo_kind":"listed","path":"src/models/canonical_heads.py","file_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners/blob/HEAD/src/models/canonical_heads.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"854be3983683cc44"}},{"code_sha256_prefix":"ae6e307cdcec33d5","entry":"measure_conflict","repo":"davidandym/task-conflict-in-text-to-text-learners","repo_kind":"listed","path":"src/measure_conflict.py","file_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners/blob/HEAD/src/measure_conflict.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ae6e307cdcec33d5"}},{"code_sha256_prefix":"e81a8138a9405edf","entry":"score","repo":"davidandym/task-conflict-in-text-to-text-learners","repo_kind":"listed","path":"src/deca_metrics.py","file_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners/blob/HEAD/src/deca_metrics.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e81a8138a9405edf"}},{"code_sha256_prefix":"7e5f9c2aee507dd4","entry":"set_seed","repo":"davidandym/task-conflict-in-text-to-text-learners","repo_kind":"listed","path":"src/utils.py","file_url":"https://github.com/davidandym/task-conflict-in-text-to-text-learners/blob/HEAD/src/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7e5f9c2aee507dd4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}