{"url":"/method/cosine-annealing","slug":"cosine-annealing","name":"Cosine Annealing","full_name":"Cosine Annealing","full_name_withheld":false,"description_markdown":"**Cosine Annealing** is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before being increased rapidly again. The resetting of the learning rate acts like a simulated restart of the learning process and the re-use of good weights as the starting point of the restart is referred to as a \"warm restart\" in contrast to a \"cold restart\" where a new set of small random numbers may be used as a starting point.\r\n\r\n$$\\eta\\_{t} = \\eta\\_{min}^{i} + \\frac{1}{2}\\left(\\eta\\_{max}^{i}-\\eta\\_{min}^{i}\\right)\\left(1+\\cos\\left(\\frac{T\\_{cur}}{T\\_{i}}\\pi\\right)\\right)\r\n$$\r\n\r\nWhere where $\\eta\\_{min}^{i}$ and $ \\eta\\_{max}^{i}$ are ranges for the learning rate, and $T\\_{cur}$ account for how many epochs have been performed since the last restart.\r\n\r\nText Source: [Jason Brownlee](https://machinelearningmastery.com/snapshot-ensemble-deep-learning-neural-network/)\r\n\r\nImage Source: [Gao Huang](https://www.researchgate.net/figure/Training-loss-of-100-layer-DenseNet-on-CIFAR10-using-standard-learning-rate-blue-and-M_fig2_315765130)","description_state":"present","introduced_year":null,"introduced_by":{"title":"SGDR: Stochastic Gradient Descent with Warm Restarts","paper":"/paper/sgdr-stochastic-gradient-descent-with-warm","first_author":"Ilya Loshchilov","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/sgdr-stochastic-gradient-descent-with-warm"},"source":{"url":"http://arxiv.org/abs/1608.03983v5","title":"SGDR: Stochastic Gradient Descent with Warm Restarts","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Learning Rate Schedules","url":"/methods/category/learning-rate-schedules","pwc_aliases":[]}],"n_papers_tagged":3965,"archive_num_papers":3965,"papers_newest_first":[{"paper":"/paper/making-language-model-a-hierarchical","title":"Making Language Model a Hierarchical Classifier and Generator","date":"2025-07-17","arxiv_id":"2507.12930","n_code_links":1,"syntology":null},{"paper":null,"title":"Generative Click-through Rate Prediction with Applications to Search Advertising","date":"2025-07-15","arxiv_id":"2507.11246","n_code_links":0,"syntology":null},{"paper":null,"title":"Behaviour Space Analysis of LLM-driven Meta-heuristic Discovery","date":"2025-07-04","arxiv_id":"2507.03605","n_code_links":0,"syntology":null},{"paper":null,"title":"Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models","date":"2025-06-28","arxiv_id":"2506.22957","n_code_links":0,"syntology":null},{"paper":null,"title":"Cat and Mouse -- Can Fake Text Generation Outpace Detector Systems?","date":"2025-06-26","arxiv_id":"2506.21274","n_code_links":0,"syntology":null},{"paper":null,"title":"Large Language Models Acing Chartered Accountancy","date":"2025-06-26","arxiv_id":"2506.21031","n_code_links":0,"syntology":null},{"paper":null,"title":"Large Language Model-Driven Code Compliance Checking in Building Information Modeling","date":"2025-06-25","arxiv_id":"2506.20551","n_code_links":0,"syntology":null},{"paper":null,"title":"Pattern-Based Phase-Separation of Tracer and Dispersed Phase Particles in Two-Phase Defocusing Particle Tracking Velocimetry","date":"2025-06-22","arxiv_id":"2506.18157","n_code_links":0,"syntology":null},{"paper":null,"title":"InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking","date":"2025-06-17","arxiv_id":"2506.14086","n_code_links":0,"syntology":null},{"paper":null,"title":"M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models","date":"2025-06-17","arxiv_id":"2506.14532","n_code_links":0,"syntology":null},{"paper":null,"title":"Toward a Graph Foundation Model: Pre-Training Transformers With Random Walks","date":"2025-06-17","arxiv_id":"2506.14098","n_code_links":0,"syntology":null},{"paper":null,"title":"Augmenting Large Language Models with Static Code Analysis for Automated Code Quality Improvements","date":"2025-06-12","arxiv_id":"2506.10330","n_code_links":0,"syntology":null},{"paper":"/paper/decomposing-mlp-activations-into","title":"Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization","date":"2025-06-12","arxiv_id":"2506.10920","n_code_links":1,"syntology":null},{"paper":"/paper/neuralnexus-at-bea-2025-shared-task-retrieval","title":"NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors","date":"2025-06-12","arxiv_id":"2506.10627","n_code_links":1,"syntology":null},{"paper":"/paper/think-before-you-simulate-symbolic-reasoning-1","title":"Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for Counterfactual Question Answering","date":"2025-06-12","arxiv_id":"2506.10753","n_code_links":1,"syntology":null},{"paper":null,"title":"A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning","date":"2025-06-11","arxiv_id":"2506.09429","n_code_links":0,"syntology":null},{"paper":null,"title":"Latent Multi-Head Attention for Small Language Models","date":"2025-06-11","arxiv_id":"2506.09342","n_code_links":0,"syntology":null},{"paper":null,"title":"AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP","date":"2025-06-10","arxiv_id":"2506.08768","n_code_links":0,"syntology":null},{"paper":"/paper/evaluating-llms-across-multi-cognitive-levels","title":"Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving","date":"2025-06-10","arxiv_id":"2506.08349","n_code_links":1,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":3}},{"paper":null,"title":"Generative Voice Bursts during Phone Call","date":"2025-06-09","arxiv_id":"2506.07526","n_code_links":0,"syntology":null},{"paper":null,"title":"LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization","date":"2025-06-09","arxiv_id":"2506.07570","n_code_links":0,"syntology":null},{"paper":null,"title":"Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models","date":"2025-06-09","arxiv_id":"2506.08147","n_code_links":0,"syntology":null},{"paper":null,"title":"Analyzing Breast Cancer Survival Disparities by Race and Demographic Location: A Survival Analysis Approach","date":"2025-06-08","arxiv_id":"2506.07191","n_code_links":0,"syntology":null},{"paper":null,"title":"Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models","date":"2025-06-08","arxiv_id":"2506.07121","n_code_links":0,"syntology":null},{"paper":null,"title":"RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation","date":"2025-06-07","arxiv_id":"2506.06677","n_code_links":0,"syntology":null},{"paper":null,"title":"Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs","date":"2025-06-06","arxiv_id":"2506.06401","n_code_links":0,"syntology":null},{"paper":null,"title":"Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey","date":"2025-06-06","arxiv_id":"2506.11102","n_code_links":0,"syntology":null},{"paper":null,"title":"The Lock-in Hypothesis: Stagnation by Algorithm","date":"2025-06-06","arxiv_id":"2506.06166","n_code_links":0,"syntology":null},{"paper":null,"title":"Benchmarking Large Language Models on Homework Assessment in Circuit Analysis","date":"2025-06-05","arxiv_id":"2506.06390","n_code_links":0,"syntology":null},{"paper":null,"title":"Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective","date":"2025-06-05","arxiv_id":"2506.05166","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":800},{"task":"/task/language-modeling","name":"Language Modeling","papers":591},{"task":"/task/question-answering","name":"Question Answering","papers":303},{"task":"/task/large-language-model","name":"Large Language Model","papers":298},{"task":"/task/text-generation","name":"Text Generation","papers":285},{"task":"/task/retrieval","name":"Retrieval","papers":203},{"task":"/task/sentence","name":"Sentence","papers":168},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":162},{"task":"/task/prompt-engineering","name":"Prompt Engineering","papers":133},{"task":"/task/few-shot-learning","name":"Few-Shot Learning","papers":111},{"task":"/task/decoder","name":"Decoder","papers":109},{"task":"/task/code-generation","name":"Code Generation","papers":106},{"task":"/task/decision-making","name":"Decision Making","papers":102},{"task":"/task/text-classification","name":"Text Classification","papers":91},{"task":"/task/rag","name":"RAG","papers":90},{"task":"/task/translation","name":"Translation","papers":89},{"task":null,"name":"GPU","papers":88},{"task":"/task/retrieval-augmented-generation","name":"Retrieval-augmented Generation","papers":88},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":88},{"task":"/task/math","name":"Math","papers":87}],"tasks_shown":20,"n_tasks":1053,"usage_by_year":[{"year":"2016","papers":2},{"year":"2018","papers":9},{"year":"2019","papers":95},{"year":"2020","papers":188},{"year":"2021","papers":308},{"year":"2022","papers":412},{"year":"2023","papers":1196},{"year":"2024","papers":1385},{"year":"2025","papers":370}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cosine-annealing"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}