{"url":"/task/l2-regularization","name":"L2 Regularization","slug":"l2-regularization","description_markdown":"See [Weight Decay](https://paperswithcode.com/method/weight-decay).\r\n\r\n**$L_{2}$ Regularization** or **Weight Decay**, is a regularization technique applied to the weights of a neural network. We minimize a loss function compromising both the primary loss function and a penalty on the $L\\_{2}$ Norm of the weights:\r\n\r\n$$L\\_{new}\\left(w\\right) = L\\_{original}\\left(w\\right) + \\lambda{w^{T}w}$$\r\n\r\nwhere $\\lambda$ is a value determining the strength of the penalty (encouraging smaller weights). \r\n\r\nWeight decay can be incorporated directly into the weight update rule, rather than just implicitly by defining it through to objective function. Often weight decay refers to the implementation where we specify it directly in the weight update rule (whereas L2 regularization is usually the implementation which is specified in the objective function).","categories":[{"name":"Methodology","url":"/area/methodology"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"derived"},"counts":{"papers_tagged":128,"papers_with_code":32,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":0,"subtasks":0,"parent_tasks":0},"benchmarks":[],"datasets":[],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":32,"tagged_in_all":128,"items":[{"url":"/paper/re-evaluating-continual-learning-scenarios-a","title":"Re-evaluating Continual Learning Scenarios: A Categorization and Case for Strong Baselines","date":"2018-10-30","arxiv_id":"1810.12488","repositories_listed":3,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/the-transient-nature-of-emergent-in-context","title":"The Transient Nature of Emergent In-Context Learning in Transformers","date":"2023-11-14","arxiv_id":"2311.08360","repositories_listed":2,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/rotational-optimizers-simple-robust-dnn","title":"Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks","date":"2023-05-26","arxiv_id":"2305.17212","repositories_listed":2,"syntology":{"n":6,"n_ran":0,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/2506-05041","title":"DACN: Dual-Attention Convolutional Network for Hyperspectral Image Super-Resolution","date":"2025-06-05","arxiv_id":"2506.05041","repositories_listed":1,"syntology":null},{"url":"/paper/walinet-a-water-and-lipid-identification","title":"WALINET: A water and lipid identification convolutional Neural Network for nuisance signal removal in 1H MR Spectroscopic Imaging","date":"2024-10-01","arxiv_id":"2410.00746","repositories_listed":1,"syntology":null},{"url":"/paper/monkeypox-disease-recognition-model-based-on","title":"Monkeypox disease recognition model based on improved SE-InceptionV3","date":"2024-03-15","arxiv_id":"2403.10087","repositories_listed":1,"syntology":null},{"url":"/paper/on-the-convergence-rate-of-the-stochastic","title":"Convergence of a L2 regularized Policy Gradient Algorithm for the Multi Armed Bandit","date":"2024-02-09","arxiv_id":"2402.06388","repositories_listed":1,"syntology":null},{"url":"/paper/prevalidated-ridge-regression-is-a-highly","title":"Prevalidated ridge regression is a highly-efficient drop-in replacement for logistic regression for high-dimensional data","date":"2024-01-28","arxiv_id":"2401.15610","repositories_listed":1,"syntology":null},{"url":"/paper/gradient-based-bilevel-optimization-for-multi","title":"Gradient-based bilevel optimization for multi-penalty Ridge regression through matrix differential calculus","date":"2023-11-23","arxiv_id":"2311.14182","repositories_listed":1,"syntology":null},{"url":"/paper/less-is-more-towards-parsimonious-multi-task","title":"Less is More -- Towards parsimonious multi-task models using structured sparsity","date":"2023-08-23","arxiv_id":"2308.12114","repositories_listed":1,"syntology":null},{"url":"/paper/maintaining-plasticity-in-deep-continual","title":"Maintaining Plasticity in Deep Continual Learning","date":"2023-06-23","arxiv_id":"2306.13812","repositories_listed":1,"syntology":{"n":5,"n_ran":1,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/it-s-enough-relaxing-diagonal-constraints-in","title":"It's Enough: Relaxing Diagonal Constraints in Linear Autoencoders for Recommendation","date":"2023-05-22","arxiv_id":"2305.12922","repositories_listed":1,"syntology":null},{"url":"/paper/planting-and-mitigating-memorized-content-in","title":"Planting and Mitigating Memorized Content in Predictive-Text Language Models","date":"2022-12-16","arxiv_id":"2212.08619","repositories_listed":1,"syntology":null},{"url":"/paper/motion-correction-and-volumetric","title":"Motion Correction and Volumetric Reconstruction for Fetal Functional Magnetic Resonance Imaging Data","date":"2022-02-11","arxiv_id":"2202.05863","repositories_listed":1,"syntology":null},{"url":"/paper/infinite-wide-finite-depth-neural-networks","title":"How Infinitely Wide Neural Networks Can Benefit from Multi-task Learning -- an Exact Macroscopic Characterization","date":"2021-12-31","arxiv_id":"2112.15577","repositories_listed":1,"syntology":null},{"url":"/paper/disturbing-target-values-for-neural-network","title":"Disturbing Target Values for Neural Network Regularization","date":"2021-10-11","arxiv_id":"2110.05003","repositories_listed":1,"syntology":null},{"url":"/paper/sequence-length-is-a-domain-length-based","title":"Sequence Length is a Domain: Length-based Overfitting in Transformer Models","date":"2021-09-15","arxiv_id":"2109.07276","repositories_listed":1,"syntology":null},{"url":"/paper/the-limitations-of-large-width-in-neural","title":"The Limitations of Large Width in Neural Networks: A Deep Gaussian Process Perspective","date":"2021-06-11","arxiv_id":"2106.06529","repositories_listed":1,"syntology":null},{"url":"/paper/learning-with-hyperspherical-uniformity","title":"Learning with Hyperspherical Uniformity","date":"2021-03-02","arxiv_id":"2103.01649","repositories_listed":1,"syntology":null},{"url":"/paper/towards-unsupervised-deep-image-enhancement","title":"Towards Unsupervised Deep Image Enhancement with Generative Adversarial Network","date":"2020-12-30","arxiv_id":"2012.15020","repositories_listed":1,"syntology":null},{"url":"/paper/neural-pruning-via-growing-regularization-1","title":"Neural Pruning via Growing Regularization","date":"2020-12-16","arxiv_id":"2012.09243","repositories_listed":1,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":3}},{"url":"/paper/label-only-membership-inference-attacks","title":"Label-Only Membership Inference Attacks","date":"2020-07-28","arxiv_id":"2007.14321","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/distributionally-robust-neural-networks","title":"Distributionally Robust Neural Networks","date":"2020-05-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/data-and-model-dependencies-of-membership","title":"Data and Model Dependencies of Membership Inference Attack","date":"2020-02-17","arxiv_id":"2002.06856","repositories_listed":1,"syntology":null},{"url":"/paper/understanding-and-stabilizing-gans-training","title":"Understanding and Stabilizing GANs' Training Dynamics with Control Theory","date":"2019-09-29","arxiv_id":"1909.13188","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/learning-a-smooth-kernel-regularizer-for","title":"Learning a smooth kernel regularizer for convolutional neural networks","date":"2019-03-05","arxiv_id":"1903.01882","repositories_listed":1,"syntology":null},{"url":"/paper/weighted-risk-minimization-deep-learning","title":"What is the Effect of Importance Weighting in Deep Learning?","date":"2018-12-08","arxiv_id":"1812.03372","repositories_listed":1,"syntology":null},{"url":"/paper/quantifying-generalization-in-reinforcement","title":"Quantifying Generalization in Reinforcement Learning","date":"2018-12-06","arxiv_id":"1812.02341","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/collaboratively-weighting-deep-and-classic","title":"Collaboratively Weighting Deep and Classic Representation via L2 Regularization for Image Classification","date":"2018-02-21","arxiv_id":"1802.07589","repositories_listed":1,"syntology":null},{"url":"/paper/convolutional-neural-networks-for-facial","title":"Convolutional Neural Networks for Facial Expression Recognition","date":"2017-04-22","arxiv_id":"1704.06756","repositories_listed":1,"syntology":null}],"syntology_records":8,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":1,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}