Methods › General › Learning Rate Schedules › Inverse Square Root Schedule
Inverse Square Root Schedule
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Inverse Square Root is a learning rate schedule 1 / √(max(n, k)) where n is the current training iteration and k is the number of warm-up steps. This sets a constant learning rate for the first k steps, then exponentially decays the learning rate until pre-training is over.
Papers archive 2025-07-28
30 shown of 702, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems 8 Jul 2025 · 0 repositories · arXiv:2507.05940
-
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution 18 Jun 2025 · 0 repositories · arXiv:2506.17323
-
Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription 17 Jun 2025 · 0 repositories · arXiv:2506.14223
-
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation 9 Jun 2025 · 0 repositories · arXiv:2506.08210
-
A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair 5 Jun 2025 · 0 repositories · arXiv:2506.04987
-
Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking 29 May 2025 · 0 repositories · arXiv:2505.23117
-
ShIOEnv: A CLI Behavior-Capturing Environment Enabling Grammar-Guided Command Synthesis for Dataset Curation 23 May 2025 · 1 repository · arXiv:2505.18374
-
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming 21 May 2025 · 0 repositories · arXiv:2505.15039Syntology ran 11 of 19 samples · 8 unverified · 19 pointer-only (licence)
-
EEG-to-Text Translation: A Model for Deciphering Human Brain Activity 20 May 2025 · 1 repository · arXiv:2505.13936
-
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation 16 May 2025 · 1 repository · arXiv:2505.11754
-
Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits 14 May 2025 · 0 repositories · arXiv:2505.09407
-
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization 8 May 2025 · 0 repositories · arXiv:2505.05070
-
GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance 7 May 2025 · 1 repository · arXiv:2505.04284
-
A review of DNA restriction-free overlapping sequence cloning techniques for synthetic biology 6 May 2025 · 0 repositories · arXiv:2505.03681
-
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry 29 Apr 2025 · 0 repositories · arXiv:2504.20849
-
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks 28 Apr 2025 · 0 repositories · arXiv:2504.19444
-
Sigma: A dataset for text-to-code semantic parsing with statistical analysis 5 Apr 2025 · 1 repository · arXiv:2504.04301
-
Advancing Sentiment Analysis in Tamil-English Code-Mixed Texts: Challenges and Transformer-Based Solutions 30 Mar 2025 · 0 repositories · arXiv:2503.23295
-
Enhancing Knowledge Graph Completion with Entity Neighborhood and Relation Context 29 Mar 2025 · 0 repositories · arXiv:2503.23205
-
Scaling Down Text Encoders of Text-to-Image Diffusion Models 25 Mar 2025 · 1 repository · arXiv:2503.19897
-
Exploring Training and Inference Scaling Laws in Generative Retrieval 24 Mar 2025 · 1 repository · arXiv:2503.18941
-
Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving 24 Mar 2025 · 0 repositories · arXiv:2503.18730
-
Enhancing Code LLM Training with Programmer Attention 19 Mar 2025 · 0 repositories · arXiv:2503.14936
-
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models 17 Mar 2025 · 0 repositories · arXiv:2503.12885
-
ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents 13 Mar 2025 · 0 repositories · arXiv:2503.10233
-
A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization 10 Mar 2025 · 0 repositories · arXiv:2503.06888
-
Roamify: Designing and Evaluating an LLM Based Google Chrome Extension for Personalised Itinerary Planning 10 Mar 2025 · 1 repository · arXiv:2504.10489
-
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform 9 Mar 2025 · 0 repositories · arXiv:2503.06676
-
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering 8 Mar 2025 · 0 repositories · arXiv:2503.06296
-
A Transformer Model for Predicting Chemical Reaction Products from Generic Templates 4 Mar 2025 · 0 repositories · arXiv:2503.05810
Tasks archive 2025-07-28
20 shown of 467 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections