{"url":"/method/exponential-decay","slug":"exponential-decay","name":"Exponential Decay","full_name":"Exponential Decay","full_name_withheld":false,"description_markdown":"**Exponential Decay** is a learning rate schedule where we decay the learning rate with more iterations using an exponential function:\r\n\r\n$$ \\text{lr} = \\text{lr}\\_{0}\\exp\\left(-kt\\right) $$\r\n\r\nImage Credit: [Suki Lau](https://towardsdatascience.com/learning-rate-schedules-and-adaptive-learning-rate-methods-for-deep-learning-2c8f433990d1)","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Learning Rate Schedules","url":"/methods/category/learning-rate-schedules","pwc_aliases":[]}],"n_papers_tagged":110,"archive_num_papers":110,"papers_newest_first":[{"paper":null,"title":"The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs","date":"2025-06-23","arxiv_id":"2506.18403","n_code_links":0,"syntology":null},{"paper":null,"title":"Compact Amplified Laser Power Stabilization Using Robust Active Disturbance Rejection Control with Sensor Noise Decoupling","date":"2025-06-10","arxiv_id":"2506.08404","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Data-Driven Model-Free Safety-Critical Control","date":"2025-06-07","arxiv_id":"2506.06931","n_code_links":0,"syntology":null},{"paper":null,"title":"Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget","date":"2025-06-03","arxiv_id":"2506.02386","n_code_links":0,"syntology":null},{"paper":null,"title":"Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models","date":"2025-05-30","arxiv_id":"2505.24187","n_code_links":0,"syntology":null},{"paper":null,"title":"Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers","date":"2025-05-27","arxiv_id":"2505.20666","n_code_links":0,"syntology":null},{"paper":null,"title":"Incremental Attractor Neural Network Modelling of the Lifespan Retrieval Curve","date":"2025-04-20","arxiv_id":"2504.14528","n_code_links":0,"syntology":null},{"paper":"/paper/on-oversquashing-in-graph-neural-networks","title":"On Oversquashing in Graph Neural Networks Through the Lens of Dynamical Systems","date":"2025-04-11","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Temporal Model On Quantum Logic","date":"2025-02-09","arxiv_id":"2502.07817","n_code_links":0,"syntology":null},{"paper":null,"title":"Uniform-in-time weak propagation of chaos for consensus-based optimization","date":"2025-02-01","arxiv_id":"2502.00582","n_code_links":0,"syntology":null},{"paper":null,"title":"Contextual Online Decision Making with Infinite-Dimensional Functional Regression","date":"2025-01-30","arxiv_id":"2501.18359","n_code_links":0,"syntology":null},{"paper":null,"title":"Hellinger-Kantorovich Gradient Flows: Global Exponential Decay of Entropy Functionals","date":"2025-01-28","arxiv_id":"2501.17049","n_code_links":0,"syntology":null},{"paper":null,"title":"The Stabilizer Bootstrap of Quantum Machine Learning with up to 10000 qubits","date":"2024-12-16","arxiv_id":"2412.11356","n_code_links":0,"syntology":null},{"paper":null,"title":"Modelling Networked Dynamical System by Temporal Graph Neural ODE with Irregularly Partial Observed Time-series Data","date":"2024-11-29","arxiv_id":"2412.00165","n_code_links":0,"syntology":null},{"paper":null,"title":"Scalable spectral representations for multi-agent reinforcement learning in network MDPs","date":"2024-10-22","arxiv_id":"2410.17221","n_code_links":0,"syntology":null},{"paper":null,"title":"Nonlinear second-order dynamics describe labial constriction trajectories across languages and contexts","date":"2024-10-10","arxiv_id":"2410.08351","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient $1$-bit tensor approximations","date":"2024-10-02","arxiv_id":"2410.01799","n_code_links":0,"syntology":null},{"paper":null,"title":"Koopman Operator in the Weighted Function Spaces and its Learning for the Estimation of Lyapunov and Zubov Functions","date":"2024-09-30","arxiv_id":"2410.00223","n_code_links":0,"syntology":null},{"paper":null,"title":"Super Level Sets and Exponential Decay: A Synergistic Approach to Stable Neural Network Training","date":"2024-09-25","arxiv_id":"2409.16769","n_code_links":0,"syntology":null},{"paper":null,"title":"Quasi-potential and drift decomposition in stochastic systems by sparse identification","date":"2024-09-10","arxiv_id":"2409.06886","n_code_links":0,"syntology":null},{"paper":null,"title":"From Few to More: Scribble-based Medical Image Segmentation via Masked Context Modeling and Continuous Pseudo Labels","date":"2024-08-23","arxiv_id":"2408.12814","n_code_links":0,"syntology":null},{"paper":null,"title":"Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models","date":"2024-08-14","arxiv_id":"2408.07472","n_code_links":0,"syntology":null},{"paper":null,"title":"Bridging Smoothness and Approximation: Theoretical Insights into Over-Smoothing in Graph Neural Networks","date":"2024-07-01","arxiv_id":"2407.01281","n_code_links":0,"syntology":null},{"paper":null,"title":"Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts","date":"2024-06-24","arxiv_id":"2406.16851","n_code_links":0,"syntology":null},{"paper":null,"title":"Finite Time Analysis of Temporal Difference Learning for Mean-Variance in a Discounted MDP","date":"2024-06-12","arxiv_id":"2406.07892","n_code_links":0,"syntology":null},{"paper":"/paper/buddy-single-channel-blind-unsupervised","title":"BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models","date":"2024-05-07","arxiv_id":"2405.04272","n_code_links":1,"syntology":null},{"paper":null,"title":"Tackling Graph Oversquashing by Global and Local Non-Dissipativity","date":"2024-05-02","arxiv_id":"2405.01009","n_code_links":0,"syntology":null},{"paper":null,"title":"Contextual Categorization Enhancement through LLMs Latent-Space","date":"2024-04-25","arxiv_id":"2404.16442","n_code_links":0,"syntology":null},{"paper":"/paper/gradformer-graph-transformer-with-exponential","title":"Gradformer: Graph Transformer with Exponential Decay","date":"2024-04-24","arxiv_id":"2404.15729","n_code_links":1,"syntology":{"ran":6,"of":9,"unverified":3,"pointer_only":9}},{"paper":null,"title":"LEMDA: A Novel Feature Engineering Method for Intrusion Detection in IoT Systems","date":"2024-04-20","arxiv_id":"2404.16870","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":12},{"task":"/task/image-classification","name":"image-classification","papers":7},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":6},{"task":"/task/classification-1","name":"Classification","papers":5},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":5},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":5},{"task":"/task/time-series-1","name":"Time Series","papers":4},{"task":"/task/time-series","name":"Time Series Analysis","papers":4},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":4},{"task":"/task/graph-neural-network","name":"Graph Neural Network","papers":3},{"task":"/task/large-language-model","name":"Large Language Model","papers":3},{"task":"/task/multi-agent-reinforcement-learning","name":"Multi-agent Reinforcement Learning","papers":3},{"task":"/task/quantum-machine-learning","name":"Quantum Machine Learning","papers":3},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":3},{"task":"/task/model","name":"model","papers":3},{"task":"/task/automl","name":"AutoML","papers":2},{"task":"/task/bayesian-inference","name":"Bayesian Inference","papers":2},{"task":"/task/change-detection","name":"Change Detection","papers":2},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":2},{"task":"/task/decision-making","name":"Decision Making","papers":2}],"tasks_shown":20,"n_tasks":107,"usage_by_year":[{"year":"2014","papers":1},{"year":"2015","papers":2},{"year":"2016","papers":2},{"year":"2017","papers":4},{"year":"2018","papers":6},{"year":"2019","papers":12},{"year":"2020","papers":10},{"year":"2021","papers":16},{"year":"2022","papers":11},{"year":"2023","papers":11},{"year":"2024","papers":23},{"year":"2025","papers":12}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/exponential-decay"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}