Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 29
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 29 of 38: papers 2,801 to 2,900 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Cost-Effective Hyperparameter Optimization for Large Language Model Generation Inference 8 Mar 2023 · 3 repositories · arXiv:2303.04673Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Stealing the Decoding Algorithms of Language Models 8 Mar 2023 · 1 repository · arXiv:2303.04729
-
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT 7 Mar 2023 · 1 repository · arXiv:2303.04226
-
Towards Zero-Shot Functional Compositionality of Language Models 6 Mar 2023 · 1 repository · arXiv:2303.03103
-
Industry Risk Assessment via Hierarchical Financial Data Using Stock Market Sentiment Indicators 5 Mar 2023 · 0 repositories · arXiv:2303.02707
-
Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners 3 Mar 2023 · 3 repositories · arXiv:2303.02151Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample)
-
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering 3 Mar 2023 · 1 repository · arXiv:2303.01903Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
WiCE: Real-World Entailment for Claims in Wikipedia 2 Mar 2023 · 2 repositories · arXiv:2303.01432Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Framework for Neurosymbolic Robot Action Planning using Large Language Models 1 Mar 2023 · 1 repository · arXiv:2303.00438Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks 1 Mar 2023 · 0 repositories · arXiv:2303.00293
-
ToxVis: Enabling Interpretability of Implicit vs. Explicit Toxicity Detection Models with Interactive Visualization 1 Mar 2023 · 0 repositories · arXiv:2303.09402
-
Zero-Shot Cross-Lingual Summarization via Large Language Models 28 Feb 2023 · 0 repositories · arXiv:2302.14229
-
Information-Restricted Neural Language Models Reveal Different Brain Regions' Sensitivity to Semantics, Syntax and Context 28 Feb 2023 · 1 repository · arXiv:2302.14389Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Large Language Models Are State-of-the-Art Evaluators of Translation Quality 28 Feb 2023 · 4 repositories · arXiv:2302.14520
-
Sampled Transformer for Point Sets 28 Feb 2023 · 0 repositories · arXiv:2302.14346
-
Inseq: An Interpretability Toolkit for Sequence Generation Models 27 Feb 2023 · 2 repositories · arXiv:2302.13942
-
LLaMA: Open and Efficient Foundation Language Models 27 Feb 2023 · 57 repositories · arXiv:2302.13971Syntology official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (of 58 harvested samples) · 4 pointer-only (licence)
-
Reward Design with Language Models 27 Feb 2023 · 1 repository · arXiv:2303.00001
-
Systematic Rectification of Language Models via Dead-end Analysis 27 Feb 2023 · 1 repository · arXiv:2302.14003Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication 26 Feb 2023 · 0 repositories · arXiv:2302.13382
-
Fast Attention Requires Bounded Entries 26 Feb 2023 · 0 repositories · arXiv:2302.13214
-
Human-in-the-Loop Schema Induction 25 Feb 2023 · 0 repositories · arXiv:2302.13048
-
Spanish Built Factual Freectianary (Spanish-BFF): the first AI-generated free dictionary 24 Feb 2023 · 0 repositories · arXiv:2302.12746
-
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning 23 Feb 2023 · 0 repositories · arXiv:2302.12313
-
What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure 23 Feb 2023 · 1 repository · arXiv:2302.12239Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
kNN-Adapter: Efficient Domain Adaptation for Black-Box Language Models 21 Feb 2023 · 0 repositories · arXiv:2302.10879
-
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey 20 Feb 2023 · 1 repository · arXiv:2302.10035
-
ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT 20 Feb 2023 · 1 repository · arXiv:2302.10205
-
A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT 18 Feb 2023 · 0 repositories · arXiv:2302.09419
-
How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation 18 Feb 2023 · 1 repository · arXiv:2302.09210Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Bounding the Capabilities of Large Language Models in Open Text Generation with Prompt Constraints 17 Feb 2023 · 1 repository · arXiv:2302.09185
-
Conveying the Predicted Future to Users: A Case Study of Story Plot Prediction 17 Feb 2023 · 1 repository · arXiv:2302.09122
-
GPT4MIA: Utilizing Generative Pre-trained Transformer (GPT-3) as A Plug-and-Play Transductive Model for Medical Image Analysis 17 Feb 2023 · 0 repositories · arXiv:2302.08722
-
PAC Prediction Sets for Large Language Models of Code 17 Feb 2023 · 1 repository · arXiv:2302.08703
-
Prompting Large Language Models With the Socratic Method 17 Feb 2023 · 0 repositories · arXiv:2303.08769
-
Foundation Models for Natural Language Processing -- Pre-trained Language Models Integrating Media 16 Feb 2023 · 0 repositories · arXiv:2302.08575
-
For Generated Text, Is NLI-Neutral Text the Best Text? 16 Feb 2023 · 1 repository · arXiv:2302.08577
-
Commonsense Reasoning for Conversational AI: A Survey of the State of the Art 15 Feb 2023 · 0 repositories · arXiv:2302.07926
-
Learning Performance-Improving Code Edits 15 Feb 2023 · 2 repositories · arXiv:2302.07867Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 11 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
Tree-Based Representation and Generation of Natural and Mathematical Language 15 Feb 2023 · 1 repository · arXiv:2302.07974
-
ScatterShot: Interactive In-context Example Curation for Text Transformation 14 Feb 2023 · 1 repository · arXiv:2302.07346
-
Diminished Diversity-of-Thought in a Standard Large Language Model 13 Feb 2023 · 0 repositories · arXiv:2302.07267
-
Can GPT-3 Perform Statutory Reasoning? 13 Feb 2023 · 1 repository · arXiv:2302.06100
-
STREET: A Multi-Task Structured Reasoning and Explanation Benchmark 13 Feb 2023 · 0 repositories · arXiv:2302.06729
-
Academic Writing with GPT-3.5: Reflections on Practices, Efficacy and Transparency 12 Feb 2023 · 0 repositories · arXiv:2304.11079
-
A Brief Report on LawGPT 1.0: A Virtual Legal Assistant Based on GPT-3 11 Feb 2023 · 0 repositories · arXiv:2302.05729
-
Combat AI With AI: Counteract Machine-Generated Fake Restaurant Reviews on Social Media 10 Feb 2023 · 1 repository · arXiv:2302.07731
-
FairPy: A Toolkit for Evaluation of Prediction Biases and their Mitigation in Large Language Models 10 Feb 2023 · 1 repository · arXiv:2302.05508
-
GTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers 10 Feb 2023 · 0 repositories · arXiv:2302.05393
-
The Wisdom of Hindsight Makes Language Models Better Instruction Followers 10 Feb 2023 · 1 repository · arXiv:2302.05206
-
Translating Natural Language to Planning Goals with Large-Language Models 10 Feb 2023 · 1 repository · arXiv:2302.05128
-
Better by you, better than me, chatgpt3 as writing assistance in students essays 9 Feb 2023 · 0 repositories · arXiv:2302.04536
-
Generating a Structured Summary of Numerous Academic Papers: Dataset and Method 9 Feb 2023 · 1 repository · arXiv:2302.04580
-
Reliable Natural Language Understanding with Large Language Models and Answer Set Programming 7 Feb 2023 · 0 repositories · arXiv:2302.03780
-
What Matters In The Structured Pruning of Generative Language Models? 7 Feb 2023 · 1 repository · arXiv:2302.03773Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples)
-
Nationality Bias in Text Generation 5 Feb 2023 · 0 repositories · arXiv:2302.02463
-
Quantized Distributed Training of Large Models with Convergence Guarantees 5 Feb 2023 · 0 repositories · arXiv:2302.02390
-
REaLTabFormer: Generating Realistic Relational and Tabular Data using Transformers 4 Feb 2023 · 3 repositories · arXiv:2302.02041Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Evaluating Large Language Models in Theory of Mind Tasks 4 Feb 2023 · 0 repositories · arXiv:2302.02083
-
Creating a Large Language Model of a Philosopher 2 Feb 2023 · 0 repositories · arXiv:2302.01339
-
Semantic Coherence Markers for the Early Diagnosis of the Alzheimer Disease 2 Feb 2023 · 1 repository · arXiv:2302.01025
-
Large language models predict human sensory judgments across six modalities 2 Feb 2023 · 0 repositories · arXiv:2302.01308
-
Analyzing Leakage of Personally Identifiable Information in Language Models 1 Feb 2023 · 1 repository · arXiv:2302.00539
-
Co-Writing with Opinionated Language Models Affects Users' Views 1 Feb 2023 · 0 repositories · arXiv:2302.00560
-
Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary Data 1 Feb 2023 · 1 repository · arXiv:2302.00674
-
An Comparative Analysis of Different Pitch and Metrical Grid Encoding Methods in the Task of Sequential Music Generation 31 Jan 2023 · 0 repositories · arXiv:2301.13383
-
Numeracy from Literacy: Data Science as an Emergent Skill from Large Language Models 31 Jan 2023 · 0 repositories · arXiv:2301.13382
-
Adaptive Machine Translation with Large Language Models 30 Jan 2023 · 1 repository · arXiv:2301.13294
-
REPLUG: Retrieval-Augmented Black-Box Language Models 30 Jan 2023 · 3 repositories · arXiv:2301.12652Syntology 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Specializing Smaller Language Models towards Multi-Step Reasoning 30 Jan 2023 · 2 repositories · arXiv:2301.12726Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A Discerning Several Thousand Judgments: GPT-3 Rates the Article + Adjective + Numeral + Noun Construction 29 Jan 2023 · 0 repositories · arXiv:2301.12564
-
Towards Equitable Representation in Text-to-Image Synthesis Models with the Cross-Cultural Understanding Benchmark (CCUB) Dataset 28 Jan 2023 · 1 repository · arXiv:2301.12073
-
The Exploration of Knowledge-Preserving Prompts for Document Summarisation 27 Jan 2023 · 0 repositories · arXiv:2301.11719
-
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning 27 Jan 2023 · 1 repository · arXiv:2301.11916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
ThoughtSource: A central hub for large language model reasoning data 27 Jan 2023 · 1 repository · arXiv:2301.11596Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Understanding the Effectiveness of Very Large Language Models on Dialog Evaluation 27 Jan 2023 · 0 repositories · arXiv:2301.12004
-
Causal Reasoning of Entities and Events in Procedural Texts 26 Jan 2023 · 1 repository · arXiv:2301.10896Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 15 harvested samples)
-
ExaRanker: Explanation-Augmented Neural Ranker 25 Jan 2023 · 1 repository · arXiv:2301.10521
-
A Stability Analysis of Fine-Tuning a Pre-Trained Model 24 Jan 2023 · 0 repositories · arXiv:2301.09820
-
Audience-Centric Natural Language Generation via Style Infusion 24 Jan 2023 · 1 repository · arXiv:2301.10283
-
The Next Chapter: A Study of Large Language Models in Storytelling 24 Jan 2023 · 0 repositories · arXiv:2301.09790
-
Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards 24 Jan 2023 · 0 repositories · arXiv:2301.10095
-
Large language models can segment narrative events similarly to humans 24 Jan 2023 · 0 repositories · arXiv:2301.10297
-
Multitask Instruction-based Prompting for Fallacy Recognition 24 Jan 2023 · 0 repositories · arXiv:2301.09992
-
AI model GPT-3 (dis)informs us better than humans 23 Jan 2023 · 0 repositories · arXiv:2301.11924
-
SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction 21 Jan 2023 · 0 repositories · arXiv:2301.08984
-
Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine 20 Jan 2023 · 1 repository · arXiv:2301.08745
-
Batch Prompting: Efficient Inference with Large Language Model APIs 19 Jan 2023 · 2 repositories · arXiv:2301.08721
-
T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations 15 Jan 2023 · 1 repository · arXiv:2301.06052Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
GPT as Knowledge Worker: A Zero-Shot Evaluation of (AI)CPA Capabilities 11 Jan 2023 · 1 repository · arXiv:2301.04408
-
Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language Models 10 Jan 2023 · 0 repositories · arXiv:2301.03797
-
Automatic Generation of German Drama Texts Using Fine Tuned GPT-2 Models 8 Jan 2023 · 0 repositories · arXiv:2301.03119
-
Critical Perspectives: A Benchmark Revealing Pitfalls in PerspectiveAPI 5 Jan 2023 · 1 repository · arXiv:2301.01874
-
Sequentially Controlled Text Generation 5 Jan 2023 · 0 repositories · arXiv:2301.02299
-
InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval 4 Jan 2023 · 1 repository · arXiv:2301.01820
-
UniHD at TSAR-2022 Shared Task: Is Compute All We Need for Lexical Simplification? 4 Jan 2023 · 1 repository · arXiv:2301.01764
-
Large Language Models as Corporate Lobbyists 3 Jan 2023 · 1 repository · arXiv:2301.01181
-
Fusing Pre-Trained Language Models With Multimodal Prompts Through Reinforcement Learning 1 Jan 2023 · 1 repository
-
Generating Human Motion From Textual Descriptions With Discrete Representations 1 Jan 2023 · 0 repositories
-
PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3 1 Jan 2023 · 0 repositories