Methods › General › Learning Rate Schedules › Cosine Annealing › Papers, page 33
Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,965 · with a code link: 1,734 · where Syntology ran a sample: 627 (513 with a run with no instrument failure, 114 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (627 of 3,965 tagged: 513 with a run with no instrument failure, 114 where every run was a failure of Syntology's instrument)
Page 33 of 40: papers 3,201 to 3,300 of 3,965, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Romantic-Computing 1 Jun 2022 · 0 repositories · arXiv:2206.11864
-
Knowledge Graph - Deep Learning: A Case Study in Question Answering in Aviation Safety Domain 31 May 2022 · 0 repositories · arXiv:2205.15952
-
Billions of Parameters Are Worth More Than In-domain Training Data: A case study in the Legal Case Entailment Task 30 May 2022 · 1 repository · arXiv:2205.15172
-
Multi-Agent Reinforcement Learning is a Sequence Modeling Problem 30 May 2022 · 1 repository · arXiv:2205.14953
-
CPED: A Large-Scale Chinese Personalized and Emotional Dialogue Dataset for Conversational AI 29 May 2022 · 1 repository · arXiv:2205.14727
-
Teaching Models to Express Their Uncertainty in Words 28 May 2022 · 1 repository · arXiv:2205.14334Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness 27 May 2022 · 13 repositories · arXiv:2205.14135Syntology official: no sample here; runs from other or unrecorded repositories · 24 ran (of which 2 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 1 violated, 17 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 30 harvested samples) · 1 pointer-only (licence)
-
kNN-Prompt: Nearest Neighbor Zero-Shot Inference 27 May 2022 · 1 repository · arXiv:2205.13792
-
What Dense Graph Do You Need for Self-Attention? 27 May 2022 · 1 repository · arXiv:2205.14014
-
Conditional set generation using Seq2seq models 25 May 2022 · 0 repositories · arXiv:2205.12485
-
Large Language Models are Few-Shot Clinical Information Extractors 25 May 2022 · 0 repositories · arXiv:2205.12689
-
NaturalProver: Grounded Mathematical Proof Generation with Language Models 25 May 2022 · 1 repository · arXiv:2205.12910Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Do we need Label Regularization to Fine-tune Pre-trained Language Models? 25 May 2022 · 0 repositories · arXiv:2205.12428
-
Transcormer: Transformer for Sentence Scoring with Sliding Language Modeling 25 May 2022 · 1 repository · arXiv:2205.12986Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
FLUTE: Figurative Language Understanding through Textual Explanations 24 May 2022 · 1 repository · arXiv:2205.12404
-
Garden-Path Traversal in GPT-2 24 May 2022 · 1 repository · arXiv:2205.12302Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
The Authenticity Gap in Human Evaluation 24 May 2022 · 0 repositories · arXiv:2205.11930
-
On the Role of Bidirectionality in Language Model Pre-Training 24 May 2022 · 0 repositories · arXiv:2205.11726
-
Improving Short Text Classification With Augmented Data Using GPT-3 23 May 2022 · 0 repositories · arXiv:2205.10981
-
Looking for a Handsome Carpenter! Debiasing GPT-3 Job Advertisements 23 May 2022 · 1 repository · arXiv:2205.11374
-
Penguins Don't Fly: Reasoning about Generics through Instantiations and Exceptions 23 May 2022 · 0 repositories · arXiv:2205.11658
-
RL with KL penalties is better viewed as Bayesian inference 23 May 2022 · 0 repositories · arXiv:2205.11275
-
GraphMAE: Self-Supervised Masked Graph Autoencoders 22 May 2022 · 3 repositories · arXiv:2205.10803Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Instruction Induction: From Few Examples to Natural Language Task Descriptions 22 May 2022 · 1 repository · arXiv:2205.10782Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models 21 May 2022 · 1 repository · arXiv:2205.10625
-
Life after BERT: What do Other Muppets Understand about Language? 21 May 2022 · 1 repository · arXiv:2205.10696
-
Prototypical Calibration for Few-shot Learning of Language Models 20 May 2022 · 1 repository · arXiv:2205.10183
-
Automated Scoring for Reading Comprehension via In-context BERT Tuning 19 May 2022 · 1 repository · arXiv:2205.09864
-
Towards Understanding Gender-Seniority Compound Bias in Natural Language Generation 19 May 2022 · 1 repository · arXiv:2205.09830
-
Evaluation of Transfer Learning for Polish with a Text-to-Text Model 18 May 2022 · 0 repositories · arXiv:2205.08808
-
M6-Rec: Generative Pretrained Language Models are Open-Ended Recommender Systems 17 May 2022 · 0 repositories · arXiv:2205.08084
-
Heroes, Villains, and Victims, and GPT-3: Automated Extraction of Character Roles Without Training Data 16 May 2022 · 0 repositories · arXiv:2205.07557
-
The AI Teacher Test: Measuring the Pedagogical Ability of Blender and GPT-3 in Educational Dialogues 16 May 2022 · 1 repository · arXiv:2205.07540
-
What GPT Knows About Who is Who 16 May 2022 · 1 repository · arXiv:2205.07407
-
Naturalistic Causal Probing for Morpho-Syntax 14 May 2022 · 1 repository · arXiv:2205.07043
-
Clinical Prompt Learning with Frozen Language Models 11 May 2022 · 1 repository · arXiv:2205.05535
-
Towards the Generation of Musical Explanations with GPT-3 11 May 2022 · 1 repository · arXiv:2206.08264
-
Object Detection in Indian Food Platters using Transfer Learning with YOLOv4 10 May 2022 · 0 repositories · arXiv:2205.04841
-
Ratatouille: A tool for Novel Recipe Generation 10 May 2022 · 0 repositories · arXiv:2206.08267
-
Reducing Activation Recomputation in Large Transformer Models 10 May 2022 · 4 repositories · arXiv:2205.05198Syntology official: harvested, nothing ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
UL2: Unifying Language Learning Paradigms 10 May 2022 · 2 repositories · arXiv:2205.05131Syntology community repositories only · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 16 harvested samples)
-
Multi-segment preserving sampling for deep manifold sampler 9 May 2022 · 0 repositories · arXiv:2205.04259
-
The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning 6 May 2022 · 1 repository · arXiv:2205.03401
-
When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it 6 May 2022 · 1 repository · arXiv:2205.03472
-
Provably Confidential Language Modelling 4 May 2022 · 1 repository · arXiv:2205.01863Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Contrastive Learning for Prompt-Based Few-Shot Language Learners 3 May 2022 · 1 repository · arXiv:2205.01308Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 6 pointer-only (licence)
-
Mixed-effects transformers for hierarchical adaptation 3 May 2022 · 1 repository · arXiv:2205.01749
-
Gradient Descent, Stochastic Optimization, and Other Tales 2 May 2022 · 0 repositories · arXiv:2205.00832
-
OPT: Open Pre-trained Transformer Language Models 2 May 2022 · 11 repositories · arXiv:2205.01068Syntology official (archive's flag): 9 ran · 14 ran (of which 2 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 17 pointer-only (licence)
-
Birds' Eye View: Measuring Behavior and Posture of Chickens as a Metric for Their Well-Being 29 Apr 2022 · 0 repositories · arXiv:2205.00069
-
Training Language Models with Language Feedback 29 Apr 2022 · 0 repositories · arXiv:2204.14146
-
Inferring Implicit Relations in Complex Questions with Language Models 28 Apr 2022 · 1 repository · arXiv:2204.13778
-
On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model 28 Apr 2022 · 0 repositories · arXiv:2204.13509
-
Tailor: A Prompt-Based Approach to Attribute-Based Controlled Text Generation 28 Apr 2022 · 0 repositories · arXiv:2204.13362
-
An End-to-End Dialogue Summarization System for Sales Calls 27 Apr 2022 · 0 repositories · arXiv:2204.12951
-
Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation 24 Apr 2022 · 1 repository · arXiv:2204.11320
-
A benchmark dataset for deep learning-based airplane detection: HRPlanes 22 Apr 2022 · 1 repository · arXiv:2204.10959
-
Measuring artificial intelligence: a systematic assessment and implications for governance 21 Apr 2022 · 0 repositories · arXiv:2204.10304
-
SinTra: Learning an inspiration model from a single multi-track music segment 21 Apr 2022 · 1 repository · arXiv:2204.09917
-
CodexDB: Generating Code for Processing SQL Queries using GPT-3 Codex 19 Apr 2022 · 0 repositories · arXiv:2204.08941
-
Impact of Tokenization on Language Models: An Analysis for Turkish 19 Apr 2022 · 0 repositories · arXiv:2204.08832
-
Zero-shot Entity and Tweet Characterization with Designed Conditional Prompts and Contexts 18 Apr 2022 · 0 repositories · arXiv:2204.08405
-
Deep Learning based Automatic Detection of Dicentric Chromosome 17 Apr 2022 · 0 repositories · arXiv:2204.08029
-
auton-survival: an Open-Source Package for Regression, Counterfactual Estimation, Evaluation and Phenotyping with Censored Time-to-Event Data 15 Apr 2022 · 3 repositories · arXiv:2204.07276
-
mGPT: Few-Shot Learners Go Multilingual 15 Apr 2022 · 1 repository · arXiv:2204.07580
-
Polling Latent Opinions: A Method for Computational Sociolinguistics Using Transformer Language Models 15 Apr 2022 · 1 repository · arXiv:2204.07483
-
Analysing similarities between legal court documents using natural language processing approaches based on Transformers 14 Apr 2022 · 0 repositories · arXiv:2204.07182
-
GPT-NeoX-20B: An Open-Source Autoregressive Language Model 14 Apr 2022 · 11 repositories · arXiv:2204.06745Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Rows from Many Sources: Enriching row completions from Wikidata with a pre-trained Language Model 14 Apr 2022 · 0 repositories · arXiv:2204.07014
-
Uniform Complexity for Text Generation 11 Apr 2022 · 1 repository · arXiv:2204.05185
-
FoundationLayerNorm: Scaling BERT and GPT to 1,000 Layers 9 Apr 2022 · 0 repositories · arXiv:2204.04477
-
Accelerating Attention through Gradient-Based Learned Runtime Pruning 7 Apr 2022 · 0 repositories · arXiv:2204.03227
-
BERTuit: Understanding Spanish language in Twitter through a native transformer 7 Apr 2022 · 0 repositories · arXiv:2204.03465
-
Testing the limits of natural language models for predicting human language judgments 7 Apr 2022 · 1 repository · arXiv:2204.03592
-
Knowledge Infused Decoding 6 Apr 2022 · 1 repository · arXiv:2204.03084Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Data Augmentation for Intent Classification with Off-the-shelf Large Language Models 5 Apr 2022 · 1 repository · arXiv:2204.01959Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems 1 Apr 2022 · 0 repositories · arXiv:2204.00212
-
Monarch: Expressive Structured Matrices for Efficient and Accurate Training 1 Apr 2022 · 2 repositories · arXiv:2204.00595Syntology community repositories only · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 1 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples)
-
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation 31 Mar 2022 · 0 repositories · arXiv:2203.16763
-
Generative Pre-Trained Transformers for Biologically Inspired Design 31 Mar 2022 · 0 repositories · arXiv:2204.09714
-
Leveraging pre-trained language models for conversational information seeking from text 31 Mar 2022 · 0 repositories · arXiv:2204.03542
-
Transformer Language Models without Positional Encodings Still Learn Positional Information 30 Mar 2022 · 1 repository · arXiv:2203.16634Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Training Compute-Optimal Large Language Models 29 Mar 2022 · 2 repositories · arXiv:2203.15556Syntology 8 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Efficient-VDVAE: Less is more 25 Mar 2022 · 1 repository · arXiv:2203.13751
-
Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory 24 Mar 2022 · 1 repository · arXiv:2203.13055Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention 23 Mar 2022 · 0 repositories · arXiv:2203.12276
-
Self-supervision through Random Segments with Autoregressive Coding (RandSAC) 22 Mar 2022 · 0 repositories · arXiv:2203.12054
-
A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots 21 Mar 2022 · 1 repository · arXiv:2203.10759
-
Compression of Generative Pre-trained Language Models via Quantization 21 Mar 2022 · 0 repositories · arXiv:2203.10705
-
Dependency-based Mixture Language Models 19 Mar 2022 · 1 repository · arXiv:2203.10256Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Analysis and Adaptation of YOLOv4 for Object Detection in Aerial Images 18 Mar 2022 · 0 repositories · arXiv:2203.10194
-
Are You Robert or RoBERTa? Deceiving Online Authorship Attribution Models Using Neural Text Generators 18 Mar 2022 · 0 repositories · arXiv:2203.09813
-
Thinking about GPT-3 In-Context Learning for Biomedical IE? Think Again 16 Mar 2022 · 1 repository · arXiv:2203.08410
-
Do Language Models Plagiarize? 15 Mar 2022 · 1 repository · arXiv:2203.07618
-
The Ghost in the Machine has an American accent: value conflict in GPT-3 15 Mar 2022 · 0 repositories · arXiv:2203.07785
-
Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations 14 Mar 2022 · 0 repositories · arXiv:2203.07511
-
GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models 14 Mar 2022 · 2 repositories · arXiv:2203.07281Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
VAST: The Valence-Assessing Semantics Test for Contextualizing Language Models 14 Mar 2022 · 1 repository · arXiv:2203.07504
-
ELLE: Efficient Lifelong Pre-training for Emerging Data 12 Mar 2022 · 1 repository · arXiv:2203.06311Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Block-Sparse Adversarial Attack to Fool Transformer-Based Text Classifiers 11 Mar 2022 · 1 repository · arXiv:2203.05948