Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 31
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 31 of 38: papers 3,001 to 3,100 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Exploring Euphemism Detection in Few-Shot and Zero-Shot Settings 24 Oct 2022 · 1 repository · arXiv:2210.12926
-
Perfectly Secure Steganography Using Minimum Entropy Coupling 24 Oct 2022 · 2 repositories · arXiv:2210.14889Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Leveraging Large Language Models for Multiple Choice Question Answering 22 Oct 2022 · 1 repository · arXiv:2210.12353Syntology official (archive's flag): 6 ran · 6 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Meta-learning Pathologies from Radiology Reports using Variance Aware Prototypical Networks 22 Oct 2022 · 0 repositories · arXiv:2210.13979
-
A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language Models 21 Oct 2022 · 1 repository · arXiv:2210.12023Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Diffuser: Efficient Transformers with Multi-hop Attention Diffusion for Long Sequences 21 Oct 2022 · 1 repository · arXiv:2210.11794
-
WikiWhy: Answering and Explaining Cause-and-Effect Questions 21 Oct 2022 · 0 repositories · arXiv:2210.12152
-
3DALL-E: Integrating Text-to-Image AI in 3D Design Workflows 20 Oct 2022 · 0 repositories · arXiv:2210.11603
-
Composing Ensembles of Pre-trained Models via Iterative Consensus 20 Oct 2022 · 0 repositories · arXiv:2210.11522
-
General Image Descriptors for Open World Image Retrieval using ViT CLIP 20 Oct 2022 · 1 repository · arXiv:2210.11141Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining 19 Oct 2022 · 4 repositories · arXiv:2210.10341
-
Towards a neural architecture of language: Deep learning versus logistics of access in neural architectures for compositional processing 19 Oct 2022 · 0 repositories · arXiv:2210.10543
-
Systematicity in GPT-3's Interpretation of Novel English Noun Compounds 18 Oct 2022 · 0 repositories · arXiv:2210.09492
-
Team Flow at DRC2022: Pipeline System for Travel Destination Recommendation Task in Spoken Dialogue 18 Oct 2022 · 0 repositories · arXiv:2210.09518
-
Tiny-Attention Adapter: Contexts Are More Important Than the Number of Parameters 18 Oct 2022 · 0 repositories · arXiv:2211.01979
-
A Generative User Simulator with GPT-based Architecture and Goal State Tracking for Reinforced Multi-Domain Dialog Systems 17 Oct 2022 · 1 repository · arXiv:2210.08692Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
Prompting GPT-3 To Be Reliable 17 Oct 2022 · 1 repository · arXiv:2210.09150
-
NormSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-Fly 16 Oct 2022 · 1 repository · arXiv:2210.08604
-
DyLoRA: Parameter Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation 14 Oct 2022 · 2 repositories · arXiv:2210.07558Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample)
-
Extracting Cultural Commonsense Knowledge at Scale 14 Oct 2022 · 2 repositories · arXiv:2210.07763
-
"John is 50 years old, can his son be 65?" Evaluating NLP Models' Understanding of Feasibility 14 Oct 2022 · 1 repository · arXiv:2210.07471
-
TestAug: A Framework for Augmenting Capability-based NLP Tests 14 Oct 2022 · 1 repository · arXiv:2210.08097
-
Saliency Map Verbalization: Comparing Feature Importance Representations from Model-free and Instruction-based Methods 13 Oct 2022 · 1 repository · arXiv:2210.07222
-
Explanations from Large Language Models Make Small Reasoners Better 13 Oct 2022 · 0 repositories · arXiv:2210.06726
-
Jointly Reinforced User Simulator and Task-oriented Dialog System with Simplified Generative Architecture 13 Oct 2022 · 0 repositories · arXiv:2210.06706
-
Language Models of Code are Few-Shot Commonsense Learners 13 Oct 2022 · 2 repositories · arXiv:2210.07128Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Large Language Models are few(1)-shot Table Reasoners 13 Oct 2022 · 1 repository · arXiv:2210.06710
-
Are Sample-Efficient NLP Models More Robust? 12 Oct 2022 · 0 repositories · arXiv:2210.06456
-
Foundation Transformers 12 Oct 2022 · 4 repositories · arXiv:2210.06423Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Predictive Querying for Autoregressive Neural Sequence Models 12 Oct 2022 · 1 repository · arXiv:2210.06464Syntology official (archive's flag): 7 ran · 7 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples)
-
SUMBot: Summarizing Context in Open-Domain Dialogue Systems 12 Oct 2022 · 0 repositories · arXiv:2210.06496
-
REV: Information-Theoretic Evaluation of Free-Text Rationales 10 Oct 2022 · 1 repository · arXiv:2210.04982Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
The Minimum Wage as an Anchor: Effects on Determinations of Fairness by Humans and AI 10 Oct 2022 · 0 repositories · arXiv:2210.10585
-
ASDOT: Any-Shot Data-to-Text Generation with Pretrained Language Models 9 Oct 2022 · 1 repository · arXiv:2210.04325
-
Controllable Dialogue Simulation with In-Context Learning 9 Oct 2022 · 1 repository · arXiv:2210.04185
-
Fine-Tuning Pre-trained Transformers into Decaying Fast Weights 9 Oct 2022 · 1 repository · arXiv:2210.04243Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
AlphaTuning: Quantization-Aware Parameter-Efficient Adaptation of Large-Scale Pre-Trained Language Models 8 Oct 2022 · 0 repositories · arXiv:2210.03858
-
Automatic Chain of Thought Prompting in Large Language Models 7 Oct 2022 · 5 repositories · arXiv:2210.03493Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
How Large Language Models are Transforming Machine-Paraphrased Plagiarism 7 Oct 2022 · 3 repositories · arXiv:2210.03568
-
Measuring and Narrowing the Compositionality Gap in Language Models 7 Oct 2022 · 1 repository · arXiv:2210.03350
-
Binding Language Models in Symbolic Languages 6 Oct 2022 · 4 repositories · arXiv:2210.02875Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Generalization Properties of Retrieval-based Models 6 Oct 2022 · 0 repositories · arXiv:2210.02617
-
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 6 Oct 2022 · 1 repository · arXiv:2210.02969
-
Rainier: Reinforced Knowledge Introspector for Commonsense Question Answering 6 Oct 2022 · 1 repository · arXiv:2210.03078Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 4 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples)
-
GLM-130B: An Open Bilingual Pre-trained Model 5 Oct 2022 · 9 repositories · arXiv:2210.02414Syntology official (archive's flag): 9 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 21 harvested samples)
-
Explaining Patterns in Data with Language Models via Interpretable Autoprompting 4 Oct 2022 · 2 repositories · arXiv:2210.01848Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Complexity-Based Prompting for Multi-Step Reasoning 3 Oct 2022 · 0 repositories · arXiv:2210.00720
-
Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought 3 Oct 2022 · 2 repositories · arXiv:2210.01240Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution 30 Sep 2022 · 2 repositories · arXiv:2210.00131
-
SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation 30 Sep 2022 · 1 repository · arXiv:2209.15323
-
Bidirectional Language Models Are Also Few-shot Learners 29 Sep 2022 · 0 repositories · arXiv:2209.14500
-
Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning 29 Sep 2022 · 2 repositories · arXiv:2209.14610Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Medical Image Captioning via Generative Pretrained Transformers 28 Sep 2022 · 0 repositories · arXiv:2209.13983
-
Who is GPT-3? An Exploration of Personality, Values and Demographics 28 Sep 2022 · 1 repository · arXiv:2209.14338
-
How GPT-3 responds to different publics on climate change and Black Lives Matter: A critical appraisal of equity in conversational AI 27 Sep 2022 · 0 repositories · arXiv:2209.13627
-
Do ever larger octopi still amplify reporting biases? Evidence from judgments of typical colour 26 Sep 2022 · 0 repositories · arXiv:2209.12786
-
News Summarization and Evaluation in the Era of GPT-3 26 Sep 2022 · 1 repository · arXiv:2209.12356
-
Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity 24 Sep 2022 · 0 repositories · arXiv:2209.12106
-
A Case Report On The "A.I. Locked-In Problem": social concerns with modern NLP 22 Sep 2022 · 0 repositories · arXiv:2209.12687
-
DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation 22 Sep 2022 · 0 repositories · arXiv:2209.10797
-
Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers 21 Sep 2022 · 0 repositories · arXiv:2209.10505
-
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering 20 Sep 2022 · 1 repository · arXiv:2209.09513Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Meta-Adapters: Parameter Efficient Few-shot Fine-tuning through Meta-Learning 19 Sep 2022 · 1 repository
-
Will It Blend? Mixing Training Paradigms & Prompting for Argument Quality Prediction 19 Sep 2022 · 0 repositories · arXiv:2209.08966
-
Psychologically-informed chain-of-thought prompts for metaphor understanding in large language models 16 Sep 2022 · 1 repository · arXiv:2209.08141Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango 16 Sep 2022 · 0 repositories · arXiv:2209.07686
-
Efficient Quantized Sparse Matrix Operations on Tensor Cores 14 Sep 2022 · 1 repository · arXiv:2209.06979
-
Out of One, Many: Using Language Models to Simulate Human Samples 14 Sep 2022 · 0 repositories · arXiv:2209.06899
-
Chain of Explanation: New Prompting Method to Generate Higher Quality Natural Language Explanation for Implicit Hate Speech 11 Sep 2022 · 0 repositories · arXiv:2209.04889
-
Why So Toxic? Measuring and Triggering Toxic Behavior in Open-Domain Chatbots 7 Sep 2022 · 0 repositories · arXiv:2209.03463
-
ChemBERTa-2: Towards Chemical Foundation Models 5 Sep 2022 · 2 repositories · arXiv:2209.01712
-
Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples 5 Sep 2022 · 0 repositories · arXiv:2209.02128
-
Do Large Language Models know what humans know? 4 Sep 2022 · 1 repository · arXiv:2209.01515
-
Every picture tells a story: Image-grounded controllable stylistic story generation 4 Sep 2022 · 0 repositories · arXiv:2209.01638
-
Elaboration-Generating Commonsense Question Answering at Scale 2 Sep 2022 · 1 repository · arXiv:2209.01232
-
FOLIO: Natural Language Reasoning with First-Order Logic 2 Sep 2022 · 1 repository · arXiv:2209.00840
-
Efficient Sparsely Activated Transformers 31 Aug 2022 · 0 repositories · arXiv:2208.14580
-
On Reality and the Limits of Language Data: Aligning LLMs with Human Norms 25 Aug 2022 · 0 repositories · arXiv:2208.11981
-
Prompting as Probing: Using Language Models for Knowledge Base Construction 23 Aug 2022 · 1 repository · arXiv:2208.11057Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
MulZDG: Multilingual Code-Switching Framework for Zero-shot Dialogue Generation 18 Aug 2022 · 1 repository · arXiv:2208.08629
-
Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies 18 Aug 2022 · 2 repositories · arXiv:2208.10264Syntology community repositories only · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 4 harvested samples)
-
Neural Embeddings for Text 17 Aug 2022 · 1 repository · arXiv:2208.08386
-
MoCapAct: A Multi-Task Dataset for Simulated Humanoid Control 15 Aug 2022 · 1 repository · arXiv:2208.07363Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Targeted Honeyword Generation with Language Models 15 Aug 2022 · 0 repositories · arXiv:2208.06946
-
Teacher Guided Training: An Efficient Framework for Knowledge Transfer 14 Aug 2022 · 0 repositories · arXiv:2208.06825
-
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models 13 Aug 2022 · 9 repositories · arXiv:2208.06677Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Debiased Large Language Models Still Associate Muslims with Uniquely Violent Acts 8 Aug 2022 · 0 repositories · arXiv:2208.04417
-
Interacting with next-phrase suggestions: How suggestion systems aid and influence the cognitive processes of writing 1 Aug 2022 · 0 repositories · arXiv:2208.00636
-
What Can Transformers Learn In-Context? A Case Study of Simple Function Classes 1 Aug 2022 · 2 repositories · arXiv:2208.01066
-
LAD: Language Models as Data for Zero-Shot Dialog 28 Jul 2022 · 0 repositories · arXiv:2207.14393
-
Large Language Models and the Reverse Turing Test 28 Jul 2022 · 0 repositories · arXiv:2207.14382
-
Is GPT-3 all you need for Visual Question Answering in Cultural Heritage? 25 Jul 2022 · 0 repositories · arXiv:2207.12101
-
Zero-Shot Video Captioning with Evolving Pseudo-Tokens 22 Jul 2022 · 1 repository · arXiv:2207.11100
-
BigIssue: A Realistic Bug Localization Benchmark 21 Jul 2022 · 0 repositories · arXiv:2207.10739
-
Word Play for Playing Othello (Reverses) 18 Jul 2022 · 0 repositories · arXiv:2207.08766
-
Can large language models reason about medical questions? 17 Jul 2022 · 1 repository · arXiv:2207.08143Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
ELECTRA is a Zero-Shot Learner, Too 17 Jul 2022 · 1 repository · arXiv:2207.08141Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Combing for Credentials: Active Pattern Extraction from Smart Reply 14 Jul 2022 · 0 repositories · arXiv:2207.10802
-
Recurrent Memory Transformer 14 Jul 2022 · 3 repositories · arXiv:2207.06881Syntology official (archive's flag): 3 ran · 6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 13 harvested samples) · 2 pointer-only (licence)
-
DynaST: Dynamic Sparse Transformer for Exemplar-Guided Image Generation 13 Jul 2022 · 1 repository · arXiv:2207.06124Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)