{"url":"/method/weight-decay","slug":"weight-decay","name":"Weight Decay","full_name":"Weight Decay","full_name_withheld":false,"description_markdown":"**Weight Decay**, or **$L_{2}$ Regularization**, is a regularization technique applied to the weights of a neural network. We minimize a loss function compromising both the primary loss function and a penalty on the $L\\_{2}$ Norm of the weights:\r\n\r\n$$L\\_{new}\\left(w\\right) = L\\_{original}\\left(w\\right) + \\lambda{w^{T}w}$$\r\n\r\nwhere $\\lambda$ is a value determining the strength of the penalty (encouraging smaller weights). \r\n\r\nWeight decay can be incorporated directly into the weight update rule, rather than just implicitly by defining it through to objective function. Often weight decay refers to the implementation where we specify it directly in the weight update rule (whereas L2 regularization is usually the implementation which is specified in the objective function).\r\n\r\nImage Source: Deep Learning, Goodfellow et al","description_state":"present","introduced_year":1943,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Regularization","url":"/methods/category/regularization","pwc_aliases":[]}],"n_papers_tagged":10713,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/constructing-and-evaluating-declarative-rag","title":"Constructing and Evaluating Declarative RAG Pipelines in PyTerrier","date":"2025-06-12","arxiv_id":"2506.10802","n_code_links":1,"syntology":null},{"paper":null,"title":"A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning","date":"2025-06-11","arxiv_id":"2506.09429","n_code_links":0,"syntology":null},{"paper":null,"title":"Auto-Compressing Networks","date":"2025-06-11","arxiv_id":"2506.09714","n_code_links":0,"syntology":null},{"paper":"/paper/learning-efficient-and-generalizable-graph","title":"Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering","date":"2025-06-11","arxiv_id":"2506.09645","n_code_links":1,"syntology":null},{"paper":"/paper/sampling-theory-for-super-resolution-with","title":"Sampling Theory for Super-Resolution with Implicit Neural Representations","date":"2025-06-11","arxiv_id":"2506.09949","n_code_links":1,"syntology":null},{"paper":null,"title":"Unsupervised Deep Clustering of MNIST with Triplet-Enhanced Convolutional Autoencoders","date":"2025-06-11","arxiv_id":"2506.10094","n_code_links":0,"syntology":null},{"paper":"/paper/llamarec-lkg-rag-a-single-pass-learnable","title":"LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking","date":"2025-06-09","arxiv_id":"2506.07449","n_code_links":1,"syntology":null},{"paper":null,"title":"LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization","date":"2025-06-09","arxiv_id":"2506.07570","n_code_links":0,"syntology":null},{"paper":null,"title":"SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding","date":"2025-06-09","arxiv_id":"2506.07600","n_code_links":0,"syntology":null},{"paper":null,"title":"Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models","date":"2025-06-08","arxiv_id":"2506.07121","n_code_links":0,"syntology":null},{"paper":null,"title":"Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs","date":"2025-06-06","arxiv_id":"2506.06401","n_code_links":0,"syntology":null},{"paper":"/paper/when-to-use-graphs-in-rag-a-comprehensive","title":"When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation","date":"2025-06-06","arxiv_id":"2506.05690","n_code_links":1,"syntology":{"ran":0,"of":5,"unverified":5,"pointer_only":0}},{"paper":null,"title":"Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning","date":"2025-06-05","arxiv_id":"2506.04527","n_code_links":0,"syntology":null},{"paper":null,"title":"On Automating Security Policies with Contemporary LLMs","date":"2025-06-05","arxiv_id":"2506.04838","n_code_links":0,"syntology":null},{"paper":null,"title":"The NTNU System at the S&I Challenge 2025 SLA Open Track","date":"2025-06-05","arxiv_id":"2506.05121","n_code_links":0,"syntology":null},{"paper":null,"title":"Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability","date":"2025-06-04","arxiv_id":"2506.03655","n_code_links":0,"syntology":null},{"paper":null,"title":"Lions and Muons: Optimization via Stochastic Frank-Wolfe","date":"2025-06-04","arxiv_id":"2506.04192","n_code_links":0,"syntology":null},{"paper":null,"title":"Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems","date":"2025-06-04","arxiv_id":"2506.03901","n_code_links":0,"syntology":null},{"paper":null,"title":"Privacy and Security Threat for OpenAI GPTs","date":"2025-06-04","arxiv_id":"2506.04036","n_code_links":0,"syntology":null},{"paper":"/paper/tracllm-a-generic-framework-for-attributing","title":"TracLLM: A Generic Framework for Attributing Long Context LLMs","date":"2025-06-04","arxiv_id":"2506.04202","n_code_links":1,"syntology":null},{"paper":null,"title":"A Novel Deep Reinforcement Learning Method for Computation Offloading in Multi-User Mobile Edge Computing with Decentralization","date":"2025-06-03","arxiv_id":"2506.02458","n_code_links":0,"syntology":null},{"paper":null,"title":"An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models","date":"2025-06-03","arxiv_id":"2506.02730","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Automatic PT Tagging for MEDLINE Citations Using Transformer-Based Models","date":"2025-06-03","arxiv_id":"2506.03321","n_code_links":0,"syntology":null},{"paper":null,"title":"Rethinking the effects of data contamination in Code Intelligence","date":"2025-06-03","arxiv_id":"2506.02791","n_code_links":0,"syntology":null},{"paper":null,"title":"LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment","date":"2025-06-02","arxiv_id":"2506.06355","n_code_links":0,"syntology":null},{"paper":null,"title":"Retrieval-Augmented Generation of Ontologies from Relational Databases","date":"2025-06-02","arxiv_id":"2506.01232","n_code_links":0,"syntology":null},{"paper":"/paper/how-neural-networks-organize-concepts","title":"How Neural Networks Organize Concepts: Introducing Concept Trajectory Analysis for Deep Learning Interpretability","date":"2025-06-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/rare-retrieval-aware-robustness-evaluation","title":"RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems","date":"2025-06-01","arxiv_id":"2506.00789","n_code_links":1,"syntology":null},{"paper":null,"title":"FinBERT2: A Specialized Bidirectional Encoder for Bridging the Gap in Finance-Specific Deployment of Large Language Models","date":"2025-05-31","arxiv_id":"2506.06335","n_code_links":0,"syntology":null},{"paper":null,"title":"Adversarial Threat Vectors and Risk Mitigation for Retrieval-Augmented Generation Systems","date":"2025-05-30","arxiv_id":"2506.00281","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":1799},{"task":"/task/language-modeling","name":"Language Modeling","papers":1406},{"task":"/task/retrieval","name":"Retrieval","papers":1368},{"task":"/task/rag","name":"RAG","papers":1244},{"task":"/task/retrieval-augmented-generation","name":"Retrieval-augmented Generation","papers":1068},{"task":"/task/question-answering","name":"Question Answering","papers":996},{"task":"/task/sentence","name":"Sentence","papers":866},{"task":"/task/large-language-model","name":"Large Language Model","papers":480},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":458},{"task":"/task/text-classification","name":"Text Classification","papers":448},{"task":"/task/text-generation","name":"Text Generation","papers":417},{"task":"/task/text-classification-1","name":"text-classification","papers":397},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":369},{"task":"/task/classification-1","name":"Classification","papers":300},{"task":"/task/information-retrieval","name":"Information Retrieval","papers":291},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":285},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":274},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":264},{"task":"/task/articles","name":"Articles","papers":255},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":254}],"tasks_shown":20,"n_tasks":1523,"usage_by_year":[{"year":"2012","papers":1},{"year":"2013","papers":3},{"year":"2014","papers":5},{"year":"2015","papers":6},{"year":"2016","papers":21},{"year":"2017","papers":21},{"year":"2018","papers":61},{"year":"2019","papers":678},{"year":"2020","papers":1425},{"year":"2021","papers":1580},{"year":"2022","papers":1211},{"year":"2023","papers":1986},{"year":"2024","papers":2658},{"year":"2025","papers":1057}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/weight-decay"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}