{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/why-do-we-need-weight-decay-in-modern-deep","title":"Why Do We Need Weight Decay in Modern Deep Learning?","arxiv_id":"2310.04415","date":"2023-10-06","proceeding":null,"authors":["Francesco D'Angelo","Maksym Andriushchenko","Aditya Varre","Nicolas Flammarion"],"abstract":"Weight decay is a broadly used technique for training state-of-the-art deep networks from image classification to large language models. Despite its widespread usage and being extensively studied in the classical literature, its role remains poorly understood for deep learning. In this work, we highlight that the role of weight decay in modern deep learning is different from its regularization effect studied in classical learning theory. For deep networks on vision tasks trained with multipass SGD, we show how weight decay modifies the optimization dynamics enhancing the ever-present implicit regularization of SGD via the loss stabilization mechanism. In contrast, for large language models trained with nearly one-epoch training, we describe how weight decay balances the bias-variance tradeoff in stochastic optimization leading to lower training loss and improved training stability. Overall, we present a unifying perspective from ResNets on vision tasks to LLMs: weight decay is never useful as an explicit regularizer but instead changes the training dynamics in a desirable way. The code is available at https://github.com/tml-epfl/why-weight-decay","url_abs":"https://arxiv.org/abs/2310.04415v2","url_pdf":"https://arxiv.org/pdf/2310.04415v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"why-do-we-need-weight-decay-in-modern-deep","repo_url":"https://github.com/tml-epfl/why-weight-decay","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"learning-theory","task_name":"Learning Theory"},{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"sgd","method_name":"SGD"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.04415","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.04415"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tml-epfl/why-weight-decay","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"e7bec5b2b1402cf8","entry":"compute_jacobian","repo":"tml-epfl/why-weight-decay","repo_kind":"official","path":"overparameterized_nets/traceh.py","file_url":"https://github.com/tml-epfl/why-weight-decay/blob/HEAD/overparameterized_nets/traceh.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e7bec5b2b1402cf8"}},{"code_sha256_prefix":"a3534f10e80ca748","entry":"load_state_dict","repo":"tml-epfl/why-weight-decay","repo_kind":"official","path":"large_language_models/model.py","file_url":"https://github.com/tml-epfl/why-weight-decay/blob/HEAD/large_language_models/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"a3534f10e80ca748"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}