Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 30
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 30 of 38: papers 2,901 to 3,000 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Rethinking with Retrieval: Faithful Large Language Model Inference 31 Dec 2022 · 1 repository · arXiv:2301.00303
-
Targeted Phishing Campaigns using Large Scale Language Models 30 Dec 2022 · 0 repositories · arXiv:2301.00665
-
GPT Takes the Bar Exam 29 Dec 2022 · 5 repositories · arXiv:2212.14402
-
Maximizing Use-Case Specificity through Precision Model Tuning 29 Dec 2022 · 0 repositories · arXiv:2212.14206
-
DeepCuts: Single-Shot Interpretability based Pruning for BERT 27 Dec 2022 · 1 repository · arXiv:2212.13392
-
TegFormer: Topic-to-Essay Generation with Good Topic Coverage and High Text Coherence 27 Dec 2022 · 0 repositories · arXiv:2212.13456
-
Using Large Language Models to Generate Engaging Captions for Data Visualizations 27 Dec 2022 · 0 repositories · arXiv:2212.14047
-
Biologically Inspired Design Concept Generation Using Generative Pre-Trained Transformers 26 Dec 2022 · 0 repositories · arXiv:2212.13196
-
Benchmark for Uncertainty & Robustness in Self-Supervised Learning 23 Dec 2022 · 1 repository · arXiv:2212.12411
-
Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? 23 Dec 2022 · 0 repositories · arXiv:2212.12131
-
Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 Surprisal 21 Dec 2022 · 1 repository · arXiv:2212.11185Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 8 pointer-only (licence)
-
JASMINE: Arabic GPT Models for Few-Shot Learning 21 Dec 2022 · 0 repositories · arXiv:2212.10755
-
KL Regularized Normalization Framework for Low Resource Tasks 21 Dec 2022 · 0 repositories · arXiv:2212.11275
-
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models 20 Dec 2022 · 1 repository · arXiv:2212.10474Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Controllable Text Generation with Language Constraints 20 Dec 2022 · 0 repositories · arXiv:2212.10466
-
Do language models have coherent mental models of everyday things? 20 Dec 2022 · 1 repository · arXiv:2212.10029Syntology official: harvested, nothing ran · 0 ran · 5 unverified (of 5 harvested samples)
-
DocAsRef: An Empirical Study on Repurposing Reference-Based Summary Quality Metrics Reference-Freely 20 Dec 2022 · 1 repository · arXiv:2212.10013
-
Generic Temporal Reasoning with Differential Analysis and Explanation 20 Dec 2022 · 0 repositories · arXiv:2212.10467
-
Go-tuning: Improving Zero-shot Learning Abilities of Smaller Language Models 20 Dec 2022 · 0 repositories · arXiv:2212.10461
-
Is GPT-3 a Good Data Annotator? 20 Dec 2022 · 1 repository · arXiv:2212.10450
-
Evaluating Psychological Safety of Large Language Models 20 Dec 2022 · 0 repositories · arXiv:2212.10529
-
Large Language Models Are Reasoning Teachers 20 Dec 2022 · 1 repository · arXiv:2212.10071Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
PairReranker: Pairwise Reranking for Natural Language Generation 20 Dec 2022 · 0 repositories · arXiv:2212.10555
-
Pay Attention to Your Tone: Introducing a New Dataset for Polite Language Rewrite 20 Dec 2022 · 1 repository · arXiv:2212.10190
-
True Detective: A Deep Abductive Reasoning Benchmark Undoable for GPT-3 and Challenging for GPT-4 20 Dec 2022 · 0 repositories · arXiv:2212.10114
-
Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers 20 Dec 2022 · 1 repository · arXiv:2212.10559
-
Emergent Analogical Reasoning in Large Language Models 19 Dec 2022 · 2 repositories · arXiv:2212.09196Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Evaluating Human-Language Model Interaction 19 Dec 2022 · 1 repository · arXiv:2212.09746
-
Large Language Models are Better Reasoners with Self-Verification 19 Dec 2022 · 1 repository · arXiv:2212.09561Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
LENS: A Learnable Evaluation Metric for Text Simplification 19 Dec 2022 · 1 repository · arXiv:2212.09739
-
Reasoning with Language Model Prompting: A Survey 19 Dec 2022 · 2 repositories · arXiv:2212.09597
-
The case for 4-bit precision: k-bit Inference Scaling Laws 19 Dec 2022 · 1 repository · arXiv:2212.09720
-
Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language Model 18 Dec 2022 · 1 repository · arXiv:2212.09146
-
MURMUR: Modular Multi-Step Reasoning for Semi-Structured Data-to-Text Generation 16 Dec 2022 · 0 repositories · arXiv:2212.08607
-
Self-Prompting Large Language Models for Zero-Shot Open-Domain QA 16 Dec 2022 · 1 repository · arXiv:2212.08635
-
Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation 15 Dec 2022 · 2 repositories · arXiv:2212.07981Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
CREPE: Can Vision-Language Foundation Models Reason Compositionally? 13 Dec 2022 · 1 repository · arXiv:2212.07796Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Paraphrase Identification with Deep Learning: A Review of Datasets and Methods 13 Dec 2022 · 0 repositories · arXiv:2212.06933
-
Elixir: Train a Large Language Model on a Small GPU Cluster 10 Dec 2022 · 2 repositories · arXiv:2212.05339
-
Thinking Fast and Slow in Large Language Models 10 Dec 2022 · 0 repositories · arXiv:2212.05206
-
Structured information extraction from complex scientific text with fine-tuned large language models 10 Dec 2022 · 0 repositories · arXiv:2212.05238
-
The Turing Deception 9 Dec 2022 · 0 repositories · arXiv:2212.06721
-
TRBLLmaker -- Transformer Reads Between Lyrics Lines maker 9 Dec 2022 · 0 repositories · arXiv:2212.04917
-
Explain to me like I am five -- Sentence Simplification Using Transformers 8 Dec 2022 · 1 repository · arXiv:2212.04595
-
LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models 8 Dec 2022 · 1 repository · arXiv:2212.04088
-
NP4G : Network Programming for Generalization 8 Dec 2022 · 1 repository · arXiv:2212.11118
-
The Role of AI in Drug Discovery: Challenges, Opportunities, and Strategies 8 Dec 2022 · 0 repositories · arXiv:2212.08104
-
DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing 7 Dec 2022 · 1 repository · arXiv:2212.03597
-
Adaptive Testing of Computer Vision Models 6 Dec 2022 · 1 repository · arXiv:2212.02774Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Counterfactual reasoning: Do language models need world knowledge for causal understanding? 6 Dec 2022 · 1 repository · arXiv:2212.03278
-
Modern French Poetry Generation with RoBERTa and GPT-2 6 Dec 2022 · 0 repositories · arXiv:2212.02911
-
Audio-Driven Co-Speech Gesture Video Generation 5 Dec 2022 · 0 repositories · arXiv:2212.02350
-
Automatic Generation of Factual News Headlines in Finnish 5 Dec 2022 · 0 repositories · arXiv:2212.02170
-
Exploring the Limits of Differentially Private Deep Learning with Group-wise Clipping 3 Dec 2022 · 0 repositories · arXiv:2212.01539
-
SumREN: Summarizing Reported Speech about Events in News 2 Dec 2022 · 1 repository · arXiv:2212.01146
-
a survey on GPT-3 1 Dec 2022 · 0 repositories · arXiv:2212.00857
-
Distilling Reasoning Capabilities into Smaller Language Models 1 Dec 2022 · 1 repository · arXiv:2212.00193
-
Quadapter: Adapter for GPT-2 Quantization 30 Nov 2022 · 0 repositories · arXiv:2211.16912
-
Outfit Generation and Recommendation -- An Experimental Study 29 Nov 2022 · 0 repositories · arXiv:2211.16353
-
Prompted Opinion Summarization with GPT-3.5 29 Nov 2022 · 1 repository · arXiv:2211.15914
-
GPT-Neo for commonsense reasoning -- a theoretical and practical lens 28 Nov 2022 · 1 repository · arXiv:2211.15593
-
Scientific and Creative Analogies in Pretrained Language Models 28 Nov 2022 · 2 repositories · arXiv:2211.15268
-
Understanding BLOOM: An empirical study on diverse NLP tasks 27 Nov 2022 · 0 repositories · arXiv:2211.14865
-
GPT-3-driven pedagogical agents for training children's curious question-asking skills 25 Nov 2022 · 0 repositories · arXiv:2211.14228
-
PromptTTS: Controllable Text-to-Speech with Text Descriptions 22 Nov 2022 · 1 repository · arXiv:2211.12171
-
Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task 21 Nov 2022 · 2 repositories · arXiv:2211.11216
-
Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification 21 Nov 2022 · 2 repositories · arXiv:2211.11158
-
PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning 21 Nov 2022 · 2 repositories · arXiv:2211.11682Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 4 pointer-only (licence)
-
Conceptor-Aided Debiasing of Large Language Models 20 Nov 2022 · 0 repositories · arXiv:2211.11087
-
A survey on knowledge-enhanced multimodal learning 19 Nov 2022 · 0 repositories · arXiv:2211.12328
-
Ignore Previous Prompt: Attack Techniques For Language Models 17 Nov 2022 · 1 repository · arXiv:2211.09527Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers 17 Nov 2022 · 1 repository · arXiv:2211.11586
-
UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization 17 Nov 2022 · 1 repository · arXiv:2211.09783
-
Galactica: A Large Language Model for Science 16 Nov 2022 · 1 repository · arXiv:2211.09085Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
TSMind: Alibaba and Soochow University's Submission to the WMT22 Translation Suggestion Task 16 Nov 2022 · 0 repositories · arXiv:2211.08987
-
GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective 15 Nov 2022 · 1 repository · arXiv:2211.08073
-
PromptCap: Prompt-Guided Task-Aware Image Captioning 15 Nov 2022 · 1 repository · arXiv:2211.09699
-
RobBERT-2022: Updating a Dutch Language Model to Account for Evolving Language Use 15 Nov 2022 · 0 repositories · arXiv:2211.08192
-
Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations 14 Nov 2022 · 1 repository · arXiv:2211.07517Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
UGIF: UI Grounded Instruction Following 14 Nov 2022 · 0 repositories · arXiv:2211.07615
-
Textual Data Augmentation for Patient Outcomes Prediction 13 Nov 2022 · 0 repositories · arXiv:2211.06778
-
Large Language Models Meet Harry Potter: A Bilingual Dataset for Aligning Dialogue Agents with Characters 13 Nov 2022 · 1 repository · arXiv:2211.06869
-
On Optimizing the Communication of Model Parallelism 10 Nov 2022 · 0 repositories · arXiv:2211.05322
-
Collateral facilitation in humans and language models 9 Nov 2022 · 1 repository · arXiv:2211.05198Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Active Example Selection for In-Context Learning 8 Nov 2022 · 1 repository · arXiv:2211.04486Syntology official (archive's flag): 10 ran · 10 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples)
-
Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device 3 Nov 2022 · 0 repositories · arXiv:2212.01217
-
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 1 Nov 2022 · 7 repositories · arXiv:2211.00593Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Text-Only Training for Image Captioning using Noise-Injected CLIP 1 Nov 2022 · 4 repositories · arXiv:2211.00575Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers 31 Oct 2022 · 17 repositories · arXiv:2210.17323Syntology official (archive's flag): 1 ran · 5 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 15 harvested samples) · 1 pointer-only (licence)
-
SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control 31 Oct 2022 · 2 repositories · arXiv:2210.17432
-
Towards Zero-Shot and Few-Shot Table Question Answering using GPT-3 31 Oct 2022 · 0 repositories · arXiv:2210.17284
-
Learning to Decompose: Hypothetical Question Decomposition Based on Comparable Texts 30 Oct 2022 · 0 repositories · arXiv:2210.16865
-
Probing for targeted syntactic knowledge through grammatical error detection 28 Oct 2022 · 1 repository · arXiv:2210.16228
-
COCO-DR: Combating Distribution Shifts in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning 27 Oct 2022 · 1 repository · arXiv:2210.15212Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
TRScore: A Novel GPT-based Readability Scorer for ASR Segmentation and Punctuation model evaluation and selection 27 Oct 2022 · 0 repositories · arXiv:2210.15104
-
Exploring Robustness of Prefix Tuning in Noisy Data: A Case Study in Financial Sentiment Analysis 26 Oct 2022 · 0 repositories · arXiv:2211.05584
-
IELM: An Open Information Extraction Benchmark for Pre-Trained Language Models 25 Oct 2022 · 0 repositories · arXiv:2210.14128
-
XRICL: Cross-lingual Retrieval-Augmented In-Context Learning for Cross-lingual Text-to-SQL Semantic Parsing 25 Oct 2022 · 0 repositories · arXiv:2210.13693
-
Inferring Past Human Actions in Homes with Abductive Reasoning 24 Oct 2022 · 1 repository · arXiv:2210.13984
-
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task 24 Oct 2022 · 4 repositories · arXiv:2210.13382Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)