Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 33
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 33 of 38: papers 3,201 to 3,300 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation 31 Mar 2022 · 0 repositories · arXiv:2203.16763
-
Generative Pre-Trained Transformers for Biologically Inspired Design 31 Mar 2022 · 0 repositories · arXiv:2204.09714
-
Leveraging pre-trained language models for conversational information seeking from text 31 Mar 2022 · 0 repositories · arXiv:2204.03542
-
Transformer Language Models without Positional Encodings Still Learn Positional Information 30 Mar 2022 · 1 repository · arXiv:2203.16634Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Training Compute-Optimal Large Language Models 29 Mar 2022 · 2 repositories · arXiv:2203.15556Syntology 8 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory 24 Mar 2022 · 1 repository · arXiv:2203.13055Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention 23 Mar 2022 · 0 repositories · arXiv:2203.12276
-
Self-supervision through Random Segments with Autoregressive Coding (RandSAC) 22 Mar 2022 · 0 repositories · arXiv:2203.12054
-
A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots 21 Mar 2022 · 1 repository · arXiv:2203.10759
-
Compression of Generative Pre-trained Language Models via Quantization 21 Mar 2022 · 0 repositories · arXiv:2203.10705
-
Dependency-based Mixture Language Models 19 Mar 2022 · 1 repository · arXiv:2203.10256Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Are You Robert or RoBERTa? Deceiving Online Authorship Attribution Models Using Neural Text Generators 18 Mar 2022 · 0 repositories · arXiv:2203.09813
-
Thinking about GPT-3 In-Context Learning for Biomedical IE? Think Again 16 Mar 2022 · 1 repository · arXiv:2203.08410
-
Do Language Models Plagiarize? 15 Mar 2022 · 1 repository · arXiv:2203.07618
-
The Ghost in the Machine has an American accent: value conflict in GPT-3 15 Mar 2022 · 0 repositories · arXiv:2203.07785
-
Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations 14 Mar 2022 · 0 repositories · arXiv:2203.07511
-
GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models 14 Mar 2022 · 2 repositories · arXiv:2203.07281Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
VAST: The Valence-Assessing Semantics Test for Contextualizing Language Models 14 Mar 2022 · 1 repository · arXiv:2203.07504
-
ELLE: Efficient Lifelong Pre-training for Emerging Data 12 Mar 2022 · 1 repository · arXiv:2203.06311Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Block-Sparse Adversarial Attack to Fool Transformer-Based Text Classifiers 11 Mar 2022 · 1 repository · arXiv:2203.05948
-
When classifying grammatical role, BERT doesn't care about word order... except when it matters 11 Mar 2022 · 1 repository · arXiv:2203.06204
-
Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction 9 Mar 2022 · 1 repository · arXiv:2203.04845Syntology official (archive's flag): 13 ran · 13 ran (of which 8 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples)
-
NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks 9 Mar 2022 · 1 repository · arXiv:2203.05081
-
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer 7 Mar 2022 · 7 repositories · arXiv:2203.03466Syntology official (archive's flag): 1 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models 4 Mar 2022 · 1 repository · arXiv:2203.02094Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Training language models to follow instructions with human feedback 4 Mar 2022 · 11 repositories · arXiv:2203.02155Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models 2 Mar 2022 · 2 repositories · arXiv:2203.01104Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Exploring and Adapting Chinese GPT to Pinyin Input Method 1 Mar 2022 · 1 repository · arXiv:2203.00249
-
A Systematic Evaluation of Large Language Models of Code 26 Feb 2022 · 3 repositories · arXiv:2202.13169Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? 25 Feb 2022 · 2 repositories · arXiv:2202.12837
-
From Natural Language to Simulations: Applying GPT-3 Codex to Automate Simulation Modeling of Logistics Systems 24 Feb 2022 · 1 repository · arXiv:2202.12107
-
Consistent Dropout for Policy Gradient Reinforcement Learning 23 Feb 2022 · 0 repositories · arXiv:2202.11818
-
SGPT: GPT Sentence Embeddings for Semantic Search 17 Feb 2022 · 1 repository · arXiv:2202.08904Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Defending against Reconstruction Attacks with Rényi Differential Privacy 15 Feb 2022 · 0 repositories · arXiv:2202.07623
-
Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam 12 Feb 2022 · 1 repository · arXiv:2202.06009
-
Semantic features of object concepts generated with GPT-3 8 Feb 2022 · 1 repository · arXiv:2202.03753
-
What are the best systems? New perspectives on NLP Benchmarking 8 Feb 2022 · 1 repository · arXiv:2202.03799
-
Cedille: A large autoregressive French language model 7 Feb 2022 · 1 repository · arXiv:2202.03371
-
Ethics, Rules of Engagement, and AI: Neural Narrative Mapping Using Large Transformer Language Models 5 Feb 2022 · 0 repositories · arXiv:2202.02647
-
A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications 4 Feb 2022 · 1 repository · arXiv:2202.02013
-
Co-training Improves Prompt-based Learning for Large Language Models 2 Feb 2022 · 1 repository · arXiv:2202.00828Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
L3Cube-MahaCorpus and MahaBERT: Marathi Monolingual Corpus, Marathi BERT Language Models, and Resources 2 Feb 2022 · 1 repository · arXiv:2202.01159
-
A Frustratingly Simple Approach for End-to-End Image Captioning 30 Jan 2022 · 0 repositories · arXiv:2201.12723
-
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models 28 Jan 2022 · 19 repositories · arXiv:2201.11903Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Describing Differences between Text Distributions with Natural Language 28 Jan 2022 · 1 repository · arXiv:2201.12323Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
DNNFuser: Generative Pre-Trained Transformer as a Generalized Mapper for Layer Fusion in DNN Accelerators 26 Jan 2022 · 0 repositories · arXiv:2201.11218
-
Synchromesh: Reliable code generation from pre-trained language models 26 Jan 2022 · 2 repositories · arXiv:2201.11227Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Pre-Trained Language Transformers are Universal Image Classifiers 25 Jan 2022 · 0 repositories · arXiv:2201.10182
-
Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection 25 Jan 2022 · 0 repositories · arXiv:2201.10474
-
Synthetic Books 24 Jan 2022 · 0 repositories · arXiv:2201.09518
-
Black-box Prompt Learning for Pre-trained Language Models 21 Jan 2022 · 1 repository · arXiv:2201.08531Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities 18 Jan 2022 · 1 repository · arXiv:2201.06796
-
A Study of Pre-trained Language Models for Analogy Generation 16 Jan 2022 · 0 repositories
-
Auto-regressive Text Generation with Pre-Trained Language Models: An Empirical Study on Question-type Short Text Generation 16 Jan 2022 · 0 repositories
-
Efficient Hierarchical Domain Adaptation for Pretrained Language Models 16 Jan 2022 · 0 repositories
-
Elastic Weight Consolidation for Reduction of Catastrophic Forgetting in GPT-2 16 Jan 2022 · 0 repositories
-
Few-Shot Semantic Parsing with Language Models Trained On Code 16 Jan 2022 · 0 repositories
-
Hierarchical Transformers Are More Efficient Language Models 16 Jan 2022 · 0 repositories
-
Jointly Reinforced User Simulator and Task-oriented Dialog System with Simplified Generative Architecture 16 Jan 2022 · 0 repositories
-
Memory-assisted prompt editing to improve GPT-3 after deployment 16 Jan 2022 · 1 repository · arXiv:2201.06009Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Penguins Don’t Fly: Reasoning about Generics through Instantiations and Exceptions 16 Jan 2022 · 0 repositories
-
Polling Latent Opinions: A Method for Computational Sociolinguistics Using Transformer Language Models 16 Jan 2022 · 0 repositories
-
Provably Confidential Language Modelling 16 Jan 2022 · 0 repositories
-
Re2G: Retrieve, Rerank, Generate 16 Jan 2022 · 0 repositories
-
Reframing Human-AI Collaboration for Generating Free-Text Explanations 16 Jan 2022 · 0 repositories
-
Representation Learning for Conversational Data using Discourse Mutual Information Maximization 16 Jan 2022 · 0 repositories
-
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models 16 Jan 2022 · 1 repository · arXiv:2201.05966Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation 16 Jan 2022 · 1 repository · arXiv:2201.05955
-
When a sentence does not introduce a discourse entity, Transformer-based models still often refer to it 16 Jan 2022 · 0 repositories
-
Why Does Surprisal From Smaller GPT-2 Models Provide Better Fit to Human Reading Times? 16 Jan 2022 · 0 repositories
-
CommonsenseQA 2.0: Exposing the Limits of AI through Gamification 14 Jan 2022 · 0 repositories · arXiv:2201.05320
-
Assemble Foundation Models for Automatic Code Summarization 13 Jan 2022 · 1 repository · arXiv:2201.05222
-
Black-Box Tuning for Language-Model-as-a-Service 10 Jan 2022 · 2 repositories · arXiv:2201.03514Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Imagined versus Remembered Stories: Quantifying Differences in Narrative Flow 7 Jan 2022 · 0 repositories · arXiv:2201.02662
-
Flow-Guided Sparse Transformer for Video Deblurring 6 Jan 2022 · 1 repository · arXiv:2201.01893Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Submix: Practical Private Prediction for Large-Scale Language Models 4 Jan 2022 · 0 repositories · arXiv:2201.00971
-
A Neural Network Solves, Explains, and Generates University Math Problems by Program Synthesis and Few-Shot Learning at Human Level 31 Dec 2021 · 1 repository · arXiv:2112.15594Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
EvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse Gate 29 Dec 2021 · 2 repositories · arXiv:2112.14397Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation 23 Dec 2021 · 3 repositories · arXiv:2112.12731
-
Few-shot Learning with Multilingual Language Models 20 Dec 2021 · 2 repositories · arXiv:2112.10668
-
Analysis and Mitigation of Dataset Artifacts in OpenAI GPT-3 19 Dec 2021 · 0 repositories
-
WebGPT: Browser-assisted question-answering with human feedback 17 Dec 2021 · 2 repositories · arXiv:2112.09332
-
Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge 16 Dec 2021 · 3 repositories · arXiv:2112.08619
-
Efficient Hierarchical Domain Adaptation for Pretrained Language Models 16 Dec 2021 · 1 repository · arXiv:2112.08786
-
Few-Shot Semantic Parsing with Language Models Trained On Code 16 Dec 2021 · 0 repositories · arXiv:2112.08696
-
Reconsidering the Past: Optimizing Hidden States in Language Models 16 Dec 2021 · 0 repositories · arXiv:2112.08653
-
Reframing Human-AI Collaboration for Generating Free-Text Explanations 16 Dec 2021 · 1 repository · arXiv:2112.08674Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Embracing Single Stride 3D Object Detector with Sparse Transformer 13 Dec 2021 · 2 repositories · arXiv:2112.06375
-
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 13 Dec 2021 · 0 repositories · arXiv:2112.06905
-
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models 13 Dec 2021 · 1 repository · arXiv:2112.06598Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Improving Logical-Level Natural Language Generation with Topic-Conditioned Data Augmentation and Logical Form Generation 12 Dec 2021 · 0 repositories · arXiv:2112.06240
-
Improving language models by retrieving from trillions of tokens 8 Dec 2021 · 2 repositories · arXiv:2112.04426Syntology 16 ran (of which 5 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 23 harvested samples) · 3 pointer-only (licence)
-
Gaudí: Conversational Interactions with Deep Representations to Generate Image Collections 5 Dec 2021 · 0 repositories · arXiv:2112.04404
-
Representation Learning for Conversational Data using Discourse Mutual Information Maximization 4 Dec 2021 · 0 repositories · arXiv:2112.05787
-
Searching for Efficient Transformers for Language Modeling 1 Dec 2021 · 0 repositories
-
Think Big, Teach Small: Do Language Models Distil Occam’s Razor? 1 Dec 2021 · 1 repository
-
Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer 1 Dec 2021 · 1 repository
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
Chemical Identification and Indexing in PubMed Articles via BERT and Text-to-Text Approaches 30 Nov 2021 · 0 repositories · arXiv:2111.15622
-
Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models 30 Nov 2021 · 1 repository · arXiv:2112.00029Syntology official: harvested, nothing ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 6 harvested samples) · 5 pointer-only (licence)