{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/broken-neural-scaling-laws","title":"Broken Neural Scaling Laws","arxiv_id":"2210.14891","date":"2022-10-26","proceeding":null,"authors":["Ethan Caballero","Kshitij Gupta","Irina Rish","David Krueger"],"abstract":"We present a smoothly broken power law functional form (referred to by us as a Broken Neural Scaling Law (BNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks (i.e. how the evaluation metric of interest varies as the amount of compute used for training, number of model parameters, training dataset size, model input size, number of training steps, or upstream performance varies) for various architectures and for each of various tasks within a large and diverse set of upstream and downstream tasks, in zero-shot, prompted, and fine-tuned settings. This set includes large-scale vision, language, audio, video, diffusion, generative modeling, multimodal learning, contrastive learning, AI alignment, robotics, out-of-distribution (OOD) generalization, continual learning, transfer learning, uncertainty estimation / calibration, out-of-distribution detection, adversarial robustness, distillation, sparsity, retrieval, quantization, pruning, molecules, computer programming/coding, math word problems, \"emergent\" \"phase transitions / changes\", arithmetic, unsupervised/self-supervised learning, and reinforcement learning (single agent and multi-agent). When compared to other functional forms for neural scaling behavior, this functional form yields extrapolations of scaling behavior that are considerably more accurate on this set. Moreover, this functional form accurately models and extrapolates scaling behavior that other functional forms are incapable of expressing such as the non-monotonic transitions present in the scaling behavior of phenomena such as double descent and the delayed, sharp inflection points present in the scaling behavior of tasks such as arithmetic. Lastly, we use this functional form to glean insights about the limit of the predictability of scaling behavior. Code is available at https://github.com/ethancaballero/broken_neural_scaling_laws","url_abs":"https://arxiv.org/abs/2210.14891v10","url_pdf":"https://arxiv.org/pdf/2210.14891v10.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"broken-neural-scaling-laws","repo_url":"https://github.com/ethancaballero/broken_neural_scaling_laws","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"adversarial-robustness","task_name":"Adversarial Robustness"},{"task_slug":"continual-learning","task_name":"Continual Learning"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"form","task_name":"Form"},{"task_slug":"math","task_name":"Math"},{"task_slug":"out-of-distribution-detection","task_name":"Out-of-Distribution Detection"},{"task_slug":"out-of-distribution-generalization","task_name":"Out-of-Distribution Generalization"},{"task_slug":"quantization","task_name":"Quantization"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.14891","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.14891"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ethancaballero/broken_neural_scaling_laws","reach":null}],"summary":{"ran_violates":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"47d0b38eff215e89","entry":"bnsl_with_1_break","repo":"ethancaballero/broken_neural_scaling_laws","repo_kind":"official","path":"fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","file_url":"https://github.com/ethancaballero/broken_neural_scaling_laws/blob/HEAD/fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"47d0b38eff215e89"}},{"code_sha256_prefix":"3353d2f8f9d44720","entry":"bnsl_with_1_break","repo":"ethancaballero/broken_neural_scaling_laws","repo_kind":"official","path":"fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","file_url":"https://github.com/ethancaballero/broken_neural_scaling_laws/blob/HEAD/fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3353d2f8f9d44720"}},{"code_sha256_prefix":"b0d4205cbf75fb83","entry":"bnsl_with_1_break__log","repo":"ethancaballero/broken_neural_scaling_laws","repo_kind":"official","path":"fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","file_url":"https://github.com/ethancaballero/broken_neural_scaling_laws/blob/HEAD/fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b0d4205cbf75fb83"}},{"code_sha256_prefix":"009a66b767de5801","entry":"bnsl_with_1_break__msle_optim","repo":"ethancaballero/broken_neural_scaling_laws","repo_kind":"official","path":"fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","file_url":"https://github.com/ethancaballero/broken_neural_scaling_laws/blob/HEAD/fit_bnsl_and_extrapolate__4_digit_addition__dataset_size_x-axis.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"009a66b767de5801"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}