Methods › General › Learning Rate Schedules › Linear Warmup
Linear Warmup
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Linear Warmup is a learning rate schedule where we linearly increase the learning rate from a low rate to a constant rate thereafter. This reduces volatility in the early stages of training.
Image Credit: Chengwei Zhang
Papers archive 2025-07-28
30 shown of 38, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence 29 May 2025 · 1 repository · arXiv:2505.23420
-
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation 20 May 2025 · 1 repository · arXiv:2505.14821
-
Tractable Representations for Convergent Approximation of Distributional HJB Equations 7 Mar 2025 · 0 repositories · arXiv:2503.05563
-
Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions 8 Jan 2025 · 0 repositories · arXiv:2501.04437
-
A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation 19 Nov 2024 · 0 repositories · arXiv:2411.12157
-
Causal Temporal Representation Learning with Nonstationary Sparse Transition 5 Sep 2024 · 1 repository · arXiv:2409.03142Syntology ran 7 of 9 samples · 2 unverified · 9 pointer-only (licence)
-
Machine learning models for daily rainfall forecasting in Northern Tropical Africa using tropical wave predictors 29 Aug 2024 · 1 repository · arXiv:2408.16349
-
CTRL: Continuous-Time Representation Learning on Temporal Heterogeneous Information Network 11 May 2024 · 0 repositories · arXiv:2405.08013
-
Towards Adversarial Robustness And Backdoor Mitigation in SSL 23 Mar 2024 · 1 repository · arXiv:2403.15918
-
Two Trades is not Baffled: Condensing Graph via Crafting Rational Gradient Matching 7 Feb 2024 · 1 repository · arXiv:2402.04924Syntology ran 3 of 3 samples · 0 unverified
-
Continual Pre-Training of Large Language Models: How to (re)warm your model? 8 Aug 2023 · 2 repositories · arXiv:2308.04014
-
CTRL: Connect Collaborative and Language Model for CTR Prediction 5 Jun 2023 · 0 repositories · arXiv:2306.02841
-
SweCTRL-Mini: a data-transparent Transformer-based large language model for controllable text generation in Swedish 27 Apr 2023 · 1 repository · arXiv:2304.13994
-
Once Detected, Never Lost: Surpassing Human Performance in Offline LiDAR based 3D Object Detection 24 Apr 2023 · 2 repositories · arXiv:2304.12315Syntology ran 3 of 5 samples · 2 unverified
-
Elastic Weight Removal for Faithful and Abstractive Dialogue Generation 30 Mar 2023 · 1 repository · arXiv:2303.17574
-
Mixing Backward- with Forward-Chaining for Metacognitive Skill Acquisition and Transfer 18 Mar 2023 · 0 repositories · arXiv:2303.12223
-
Alternative formulations for gilthead seabream diets: towards a more sustainable production 3 Nov 2022 · 0 repositories · arXiv:2211.02430
-
Controllable Factuality in Document-Grounded Dialog Systems Using a Noisy Channel Model 31 Oct 2022 · 1 repository · arXiv:2210.17418
-
Unsupervised Learning of Structured Representations via Closed-Loop Transcription 30 Oct 2022 · 1 repository · arXiv:2210.16782
-
An Embarrassingly Simple Backdoor Attack on Self-supervised Learning 13 Oct 2022 · 4 repositories · arXiv:2210.07346Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
Continual Learning, Fast and Slow 6 Sep 2022 · 1 repository · arXiv:2209.02370
-
CTRL: Clustering Training Losses for Label Error Detection 17 Aug 2022 · 1 repository · arXiv:2208.08464
-
Pursuit of a Discriminative Representation for Multiple Subspaces via Sequential Games 18 Jun 2022 · 1 repository · arXiv:2206.09120Syntology ran 0 of 1 samples · 1 unverified
-
Efficient and Training-Free Control of Language Generation 12 May 2022 · 0 repositories · arXiv:2205.06036
-
Gradient Descent, Stochastic Optimization, and Other Tales 2 May 2022 · 0 repositories · arXiv:2205.00832
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation 27 Sep 2021 · 3 repositories · arXiv:2109.13296Syntology ran 0 of 7 samples · 7 unverified
-
Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL 4 Jun 2021 · 1 repository · arXiv:2106.02193Syntology ran 7 of 11 samples · 4 unverified · 11 pointer-only (licence)
-
Learning to Relate Depth and Semantics for Unsupervised Domain Adaptation 17 May 2021 · 1 repository · arXiv:2105.07830Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)
-
CTLR@WiC-TSV: Target Sense Verification using Marked Inputs andPre-trained Models 30 Apr 2021 · 0 repositories
Tasks archive 2025-07-28
20 shown of 79 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections