Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 34
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 34 of 38: papers 3,301 to 3,400 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Context Matters in Semantically Controlled Language Generation for Task-oriented Dialogue Systems 28 Nov 2021 · 0 repositories · arXiv:2111.14119
-
Domain Prompt Learning for Efficiently Adapting CLIP to Unseen Domains 25 Nov 2021 · 1 repository · arXiv:2111.12853Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Transformer-based Korean Pretrained Language Models: A Survey on Three Years of Progress 25 Nov 2021 · 0 repositories · arXiv:2112.03014
-
ClipCap: CLIP Prefix for Image Captioning 18 Nov 2021 · 4 repositories · arXiv:2111.09734Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN 18 Nov 2021 · 0 repositories · arXiv:2111.09509
-
Guiding Generative Language Models for Data Augmentation in Few-Shot Text Classification 17 Nov 2021 · 0 repositories · arXiv:2111.09064
-
Active Dialogue Simulation in Conversational Systems 16 Nov 2021 · 0 repositories
-
An Information Theoretic Measurement of Topical Relevance in Learner Essays 16 Nov 2021 · 0 repositories
-
Building Chinese Biomedical Language Models via Multi-Level Text Discrimination 16 Nov 2021 · 1 repository
-
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation 16 Nov 2021 · 0 repositories
-
Data Augmentation for Intent Classification with Generic Large Language Models 16 Nov 2021 · 0 repositories
-
ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language Models 16 Nov 2021 · 0 repositories
-
ELLE: Efficient Lifelong Pre-training for Emerging Data 16 Nov 2021 · 0 repositories
-
End-to-end Task-oriented Dialog Policy Learning based on Pre-trained Language Model 16 Nov 2021 · 0 repositories
-
ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention 16 Nov 2021 · 0 repositories
-
Exploring and Adapting Chinese GPT to Pinyin Input Method 16 Nov 2021 · 0 repositories
-
Generative Pre-Trained Transformer for Design Concept Generation: An Exploration 16 Nov 2021 · 0 repositories · arXiv:2111.08489
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 16 Nov 2021 · 0 repositories
-
Impact of Tokenization on Language Models: An Analysis for Turkish 16 Nov 2021 · 0 repositories
-
Improving GPT-3 after deployment with a dynamic memory of feedback 16 Nov 2021 · 0 repositories
-
Knowledge Graph is in Rescue: Task Oriented Dialogue System for Response Generation without NLU and DM 16 Nov 2021 · 0 repositories
-
Life after BERT: What do Other Muppets Understand about Language? 16 Nov 2021 · 0 repositories
-
Moving the Eiffel Tower to ROME: Tracing and Editing Facts in GPT 16 Nov 2021 · 0 repositories
-
On the Multilingual Capabilities of Very Large-Scale English Language Models 16 Nov 2021 · 0 repositories
-
Representation of Ambiguity in Pre-Trained Sentence Embeddings 16 Nov 2021 · 0 repositories
-
Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word Distributions 16 Nov 2021 · 0 repositories
-
Tell me who you are and i'll tell you what to do: A Persona Grounded Task Oriented Dialogue Generation System 16 Nov 2021 · 0 repositories
-
The Power of Prompt Tuning for Low-Resource Semantic Parsing 16 Nov 2021 · 0 repositories
-
Towards Coding Social Science Datasets with Language Models 16 Nov 2021 · 0 repositories
-
When classifying grammatical role, BERT doesn't care about word order... except when it matters 16 Nov 2021 · 0 repositories
-
Exploring Story Generation with Multi-task Objectives in Variational Autoencoders 15 Nov 2021 · 0 repositories · arXiv:2111.08133
-
Scaling Law for Recommendation Models: Towards General-purpose User Representations 15 Nov 2021 · 0 repositories · arXiv:2111.11294
-
A Novel Corpus of Discourse Structure in Humans and Computers 10 Nov 2021 · 1 repository · arXiv:2111.05940
-
Amazon SageMaker Model Parallelism: A General and Flexible Framework for Large Model Training 10 Nov 2021 · 0 repositories · arXiv:2111.05972
-
DistIR: An Intermediate Representation and Simulator for Efficient Neural Network Distribution 9 Nov 2021 · 0 repositories · arXiv:2111.05426
-
FPM: A Collection of Large-scale Foundation Pre-trained Language Models 9 Nov 2021 · 0 repositories · arXiv:2111.04909
-
TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches 8 Nov 2021 · 2 repositories · arXiv:2111.04867
-
An Explanation of In-context Learning as Implicit Bayesian Inference 3 Nov 2021 · 1 repository · arXiv:2111.02080Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
TheEyeCorpus: Experiments in Reducing NLP Bias and Identifiability for Large LMs 3 Nov 2021 · 0 repositories
-
DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language Models 30 Oct 2021 · 1 repository · arXiv:2111.00160Syntology official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 10 pointer-only (licence)
-
Amendable Generation for Dialogue State Tracking 29 Oct 2021 · 1 repository · arXiv:2110.15659
-
A Sequence to Sequence Model for Extracting Multiple Product Name Entities from Dialog 28 Oct 2021 · 0 repositories · arXiv:2110.14843
-
Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training 28 Oct 2021 · 1 repository · arXiv:2110.14883
-
Generating artificial texts as substitution or complement of training data 25 Oct 2021 · 0 repositories · arXiv:2110.13016
-
Fast Model Editing at Scale 21 Oct 2021 · 3 repositories · arXiv:2110.11309Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Risks of AI Foundation Models in Education 19 Oct 2021 · 0 repositories · arXiv:2110.10024
-
Illiterate DALL-E Learns to Compose 17 Oct 2021 · 1 repository · arXiv:2110.11405Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Reminding the Incremental Language Model via Data-Free Self-Distillation 17 Oct 2021 · 0 repositories · arXiv:2110.08745
-
Taming Visually Guided Sound Generation 17 Oct 2021 · 3 repositories · arXiv:2110.08791Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
A Short Study on Compressing Decoder-Based Language Models 16 Oct 2021 · 0 repositories · arXiv:2110.08460
-
Evaluation of Transfer Learning for Polish with a text-to-text model 16 Oct 2021 · 0 repositories
-
Hydra: A System for Large Multi-Model Deep Learning 16 Oct 2021 · 1 repository · arXiv:2110.08633
-
Knowledge Inheritance for Pre-trained Language Models 16 Oct 2021 · 0 repositories
-
PAGnol: An Extra-Large French Generative Model 16 Oct 2021 · 0 repositories · arXiv:2110.08554
-
Sharpness-Aware Minimization Improves Language Model Generalization 16 Oct 2021 · 0 repositories · arXiv:2110.08529
-
The Power of Prompt Tuning for Low-Resource Semantic Parsing 16 Oct 2021 · 0 repositories · arXiv:2110.08525
-
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models 16 Oct 2021 · 0 repositories
-
Kronecker Decomposition for GPT Compression 15 Oct 2021 · 0 repositories · arXiv:2110.08152
-
Building Chinese Biomedical Language Models via Multi-Level Text Discrimination 14 Oct 2021 · 1 repository · arXiv:2110.07244
-
Can Machines Learn Morality? The Delphi Experiment 14 Oct 2021 · 1 repository · arXiv:2110.07574
-
Symbolic Knowledge Distillation: from General Language Models to Commonsense Models 14 Oct 2021 · 1 repository · arXiv:2110.07178
-
Leveraging Generative Models for Covert Messaging: Challenges and Tradeoffs for "Dead-Drop" Deployments 13 Oct 2021 · 0 repositories · arXiv:2110.07009
-
Language Modelling via Learning to Rank 13 Oct 2021 · 0 repositories · arXiv:2110.06961
-
Scaling Laws for the Few-Shot Adaptation of Pre-trained Image Classifiers 13 Oct 2021 · 0 repositories · arXiv:2110.06990
-
Yformer: U-Net Inspired Transformer Architecture for Far Horizon Time Series Forecasting 13 Oct 2021 · 1 repository · arXiv:2110.08255
-
LightSeq2: Accelerated Training for Transformer-based Models on GPUs 12 Oct 2021 · 1 repository · arXiv:2110.05722
-
LiST: Lite Prompted Self-training Makes Parameter-Efficient Few-shot Learners 12 Oct 2021 · 1 repository · arXiv:2110.06274Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 2 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 4 pointer-only (licence)
-
Multi-Task Learning for Situated Multi-Domain End-to-End Dialogue Systems 11 Oct 2021 · 0 repositories · arXiv:2110.05221
-
DCT: Dynamic Compressive Transformer for Modeling Unbounded Sequence 10 Oct 2021 · 0 repositories · arXiv:2110.04821
-
Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning 10 Oct 2021 · 1 repository · arXiv:2110.04725
-
Vector-quantized Image Modeling with Improved VQGAN 9 Oct 2021 · 5 repositories · arXiv:2110.04627Syntology 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining 8 Oct 2021 · 0 repositories · arXiv:2110.03888
-
Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling 7 Oct 2021 · 1 repository · arXiv:2110.03252
-
Leveraging the Inductive Bias of Large Language Models for Abstract Textual Reasoning 5 Oct 2021 · 0 repositories · arXiv:2110.02370
-
Word Acquisition in Neural Language Models 5 Oct 2021 · 1 repository · arXiv:2110.02406
-
Adversarial Examples Generation for Reducing Implicit Gender Bias in Pre-trained Models 3 Oct 2021 · 0 repositories · arXiv:2110.01094
-
Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models 1 Oct 2021 · 0 repositories · arXiv:2110.00672
-
Collaborative Storytelling with Human Actors and AI Narrators 29 Sep 2021 · 0 repositories · arXiv:2109.14728
-
ERNIE-SPARSE: Robust Efficient Transformer Through Hierarchically Unifying Isolated Information 29 Sep 2021 · 0 repositories
-
Illiterate DALL·E Learns to Compose 29 Sep 2021 · 0 repositories
-
Language Model Pre-training Improves Generalization in Policy Learning 29 Sep 2021 · 0 repositories
-
Mapping Language Models to Grounded Conceptual Spaces 29 Sep 2021 · 0 repositories
-
Offline Reinforcement Learning for Large Scale Language Action Spaces 29 Sep 2021 · 0 repositories
-
SeqPATE: Differentially Private Text Generation via Knowledge Distillation 29 Sep 2021 · 0 repositories
-
RAFT: A Real-World Few-Shot Text Classification Benchmark 28 Sep 2021 · 1 repository · arXiv:2109.14076
-
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation 27 Sep 2021 · 3 repositories · arXiv:2109.13296Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Language Models as Recommender Systems: Evaluations and Limitations 22 Sep 2021 · 0 repositories
-
Recursively Summarizing Books with Human Feedback 22 Sep 2021 · 0 repositories · arXiv:2109.10862
-
A Plug-and-Play Method for Controlled Text Generation 20 Sep 2021 · 1 repository · arXiv:2109.09707
-
Model Bias in NLP -- Application to Hate Speech Classification using transfer learning techniques 20 Sep 2021 · 0 repositories · arXiv:2109.09725
-
Towards Zero-Label Language Learning 19 Sep 2021 · 0 repositories · arXiv:2109.09193
-
Learning Low-frequency Patterns with A Pre-trained Document-Grounded Conversation Model 17 Sep 2021 · 0 repositories
-
Primer: Searching for Efficient Transformers for Language Modeling 17 Sep 2021 · 4 repositories · arXiv:2109.08668Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Relating Neural Text Degeneration to Exposure Bias 17 Sep 2021 · 0 repositories · arXiv:2109.08705
-
Language Models are Few-shot Multilingual Learners 16 Sep 2021 · 1 repository · arXiv:2109.07684Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Improving Text Auto-Completion with Next Phrase Prediction 15 Sep 2021 · 0 repositories · arXiv:2109.07067
-
A Temporal Variational Model for Story Generation 14 Sep 2021 · 3 repositories · arXiv:2109.06807
-
Multilingual Translation via Grafting Pre-trained Language Models 11 Sep 2021 · 1 repository · arXiv:2109.05256
-
TopicRefine: Joint Topic Prediction and Dialogue Response Generation for Multi-turn End-to-End Dialogue System 11 Sep 2021 · 0 repositories · arXiv:2109.05187
-
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA 10 Sep 2021 · 1 repository · arXiv:2109.05014Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)