Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers, page 36
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Page 36 of 38: papers 3,501 to 3,600 of 3,797, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
KI-BERT: Infusing Knowledge Context for Better Language and Domain Understanding 9 Apr 2021 · 0 repositories · arXiv:2104.08145
-
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking 9 Apr 2021 · 1 repository · arXiv:2104.04466
-
Using GPT-2 to Create Synthetic Data to Improve the Prediction Performance of NLP Machine Learning Classification Models 2 Apr 2021 · 0 repositories · arXiv:2104.10658
-
Russian Paraphrasers: Paraphrase with Transformers 1 Apr 2021 · 2 repositories
-
Automatic Graph Partitioning for Very Large-scale Deep Learning 30 Mar 2021 · 0 repositories · arXiv:2103.16063
-
FastMoE: A Fast Mixture-of-Expert Training System 24 Mar 2021 · 3 repositories · arXiv:2103.13262Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Thinking Aloud: Dynamic Context Generation Improves Zero-Shot Reasoning Performance of GPT-2 24 Mar 2021 · 0 repositories · arXiv:2103.13033
-
Detecting Hate Speech with GPT-3 23 Mar 2021 · 2 repositories · arXiv:2103.12407
-
The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures 23 Mar 2021 · 0 repositories · arXiv:2104.10640
-
Play the Shannon Game With Language Models: A Human-Free Approach to Summary Evaluation 19 Mar 2021 · 0 repositories · arXiv:2103.10918
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 18 Mar 2021 · 8 repositories · arXiv:2103.10360Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
GPT Understands, Too 18 Mar 2021 · 10 repositories · arXiv:2103.10385Syntology official (archive's flag): 1 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Large Pre-trained Language Models Contain Human-like Biases of What is Right and Wrong to Do 8 Mar 2021 · 1 repository · arXiv:2103.11790Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Disentangling Syntax and Semantics in the Brain with Deep Networks 2 Mar 2021 · 0 repositories · arXiv:2103.01620
-
Long Document Summarization in a Low Resource Setting using Pretrained Language Models 1 Mar 2021 · 0 repositories · arXiv:2103.00751
-
From Universal Language Model to Downstream Task: Improving RoBERTa-Based Vietnamese Hate Speech Detection 24 Feb 2021 · 0 repositories · arXiv:2102.12162
-
Robust and Transferable Anomaly Detection in Log Data using Pre-Trained Language Models 23 Feb 2021 · 0 repositories · arXiv:2102.11570
-
Calibrate Before Use: Improving Few-Shot Performance of Language Models 19 Feb 2021 · 5 repositories · arXiv:2102.09690Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
THEaiTRE 1.0: Interactive generation of theatre play scripts 17 Feb 2021 · 0 repositories · arXiv:2102.08892
-
Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet 16 Feb 2021 · 1 repository · arXiv:2102.08036
-
TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models 16 Feb 2021 · 1 repository · arXiv:2102.07988Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm 15 Feb 2021 · 0 repositories · arXiv:2102.07350
-
The corruptive force of AI-generated advice 15 Feb 2021 · 0 repositories · arXiv:2102.07536
-
Multiversal views on language models 12 Feb 2021 · 0 repositories · arXiv:2102.06391
-
AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models 9 Feb 2021 · 1 repository · arXiv:2102.05126
-
A Hybrid Task-Oriented Dialog System with Domain and Task Adaptive Pretraining 8 Feb 2021 · 0 repositories · arXiv:2102.04506
-
Generating Fake Cyber Threat Intelligence Using Transformer-Based Models 8 Feb 2021 · 0 repositories · arXiv:2102.04351
-
Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models 8 Feb 2021 · 1 repository · arXiv:2102.04130Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic Labeling 6 Feb 2021 · 0 repositories · arXiv:2102.03551
-
Neural Data-to-Text Generation with LM-based Text Augmentation 6 Feb 2021 · 0 repositories · arXiv:2102.03556
-
PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers 5 Feb 2021 · 1 repository · arXiv:2102.03161Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Understanding Emails and Drafting Responses -- An Approach Using GPT-3 5 Feb 2021 · 0 repositories · arXiv:2102.03062
-
Adaptive Semiparametric Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02557
-
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02503
-
Mind the Gap: Assessing Temporal Generalization in Neural Language Models 3 Feb 2021 · 1 repository · arXiv:2102.01951
-
"Is depression related to cannabis?": A knowledge-infused model for Entity and Relation Extraction with Limited Supervision 1 Feb 2021 · 0 repositories · arXiv:2102.01222
-
Synthesizing Monolingual Data for Neural Machine Translation 29 Jan 2021 · 0 repositories · arXiv:2101.12462
-
BERT Transformer model for Detecting Arabic GPT2 Auto-Generated Tweets 22 Jan 2021 · 0 repositories · arXiv:2101.09345
-
Towards Facilitating Empathic Conversations in Online Mental Health Support: A Reinforcement Learning Approach 19 Jan 2021 · 1 repository · arXiv:2101.07714
-
Persistent Anti-Muslim Bias in Large Language Models 14 Jan 2021 · 1 repository · arXiv:2101.05783
-
Adding Recurrence to Pretrained Transformers 1 Jan 2021 · 0 repositories
-
Cluster-Former: Clustering-based Sparse Transformer for Question Answering 1 Jan 2021 · 0 repositories
-
How Multipurpose Are Language Models? 1 Jan 2021 · 0 repositories
-
KETG: A Knowledge Enhanced Text Generation Framework 1 Jan 2021 · 0 repositories
-
Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models 1 Jan 2021 · 1 repository · arXiv:2101.00288
-
Prefix-Tuning: Optimizing Continuous Prompts for Generation 1 Jan 2021 · 13 repositories · arXiv:2101.00190Syntology community repositories only · 4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Pretrain Knowledge-Aware Language Models 1 Jan 2021 · 0 repositories
-
Subformer: A Parameter Reduced Transformer 1 Jan 2021 · 0 repositories
-
VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words 1 Jan 2021 · 1 repository · arXiv:2101.00265
-
WARP: Word-level Adversarial ReProgramming 1 Jan 2021 · 1 repository · arXiv:2101.00121
-
Conditional Generation of Temporally-ordered Event Sequences 31 Dec 2020 · 0 repositories · arXiv:2012.15786
-
Directed Beam Search: Plug-and-Play Lexically Constrained Language Generation 31 Dec 2020 · 1 repository · arXiv:2012.15416Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Making Pre-trained Language Models Better Few-shot Learners 31 Dec 2020 · 9 repositories · arXiv:2012.15723Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 9 harvested samples) · 7 pointer-only (licence)
-
The Pile: An 800GB Dataset of Diverse Text for Language Modeling 31 Dec 2020 · 22 repositories · arXiv:2101.00027Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Robust Dialogue Utterance Rewriting as Sequence Tagging 29 Dec 2020 · 1 repository · arXiv:2012.14535
-
Uncertainty and Surprisal Jointly Deliver the Punchline: Exploiting Incongruity-Based Features for Humor Recognition 22 Dec 2020 · 0 repositories · arXiv:2012.12007
-
Breaking Writer's Block: Low-cost Fine-tuning of Natural Language Generation Models 19 Dec 2020 · 0 repositories · arXiv:2101.03216
-
Query expansion with artificially generated texts 16 Dec 2020 · 0 repositories · arXiv:2012.08787
-
Revisiting Linformer with a modified self-attention with linear complexity 16 Dec 2020 · 0 repositories · arXiv:2101.10277
-
RecipeNLG: A Cooking Recipes Dataset for Semi-Structured Text Generation 15 Dec 2020 · 1 repository
-
Extracting Training Data from Large Language Models 14 Dec 2020 · 3 repositories · arXiv:2012.07805Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment 11 Dec 2020 · 0 repositories · arXiv:2012.06373
-
As Good as New. How to Successfully Recycle English GPT-2 to Make Models for Other Languages 10 Dec 2020 · 1 repository · arXiv:2012.05628
-
Towards Neural Programming Interfaces 10 Dec 2020 · 1 repository · arXiv:2012.05983
-
CX DB8: A queryable extractive summarizer and semantic search engine 7 Dec 2020 · 2 repositories · arXiv:2012.03942
-
UBAR: Towards Fully End-to-End Task-Oriented Dialog Systems with GPT-2 7 Dec 2020 · 1 repository · arXiv:2012.03539
-
Enhanced Offensive Language Detection Through Data Augmentation 5 Dec 2020 · 0 repositories · arXiv:2012.02954
-
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 4 Dec 2020 · 0 repositories · arXiv:2012.02469
-
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering 2 Dec 2020 · 1 repository · arXiv:2012.00955
-
A Deep Generative Approach to Native Language Identification 1 Dec 2020 · 0 repositories
-
Adversarial Sparse Transformer for Time Series Forecasting 1 Dec 2020 · 1 repository
-
Comparing Probabilistic, Distributional and Transformer-Based Models on Logical Metonymy Interpretation 1 Dec 2020 · 0 repositories
-
CPM: A Large-scale Generative Chinese Pre-trained Language Model 1 Dec 2020 · 10 repositories · arXiv:2012.00413
-
Hitachi at SemEval-2020 Task 11: An Empirical Study of Pre-Trained Transformer Family for Propaganda Detection 1 Dec 2020 · 0 repositories
-
Hitachi at SemEval-2020 Task 7: Stacking at Scale with Heterogeneous Language Models for Humor Recognition 1 Dec 2020 · 0 repositories
-
Hitachi at SemEval-2020 Task 8: Simple but Effective Modality Ensemble for Meme Emotion Recognition 1 Dec 2020 · 0 repositories
-
Increasing Learning Efficiency of Self-Attention Networks through Direct Position Interactions, Learnable Temperature, and Convoluted Attention 1 Dec 2020 · 1 repository
-
TableGPT: Few-shot Table-to-Text Generation with Table Structure Reconstruction and Content Matching 1 Dec 2020 · 1 repository
-
UI at SemEval-2020 Task 4: Commonsense Validation and Explanation by Exploiting Contradiction 1 Dec 2020 · 0 repositories
-
Generative Pre-training for Paraphrase Generation by Representing and Predicting Spans in Exemplars 29 Nov 2020 · 0 repositories · arXiv:2011.14344
-
An Efficient and Scalable Deep Learning Approach for Road Damage Detection 18 Nov 2020 · 2 repositories · arXiv:2011.09577
-
Do Fine-tuned Commonsense Language Models Really Generalize? 18 Nov 2020 · 0 repositories · arXiv:2011.09159
-
Attention Mechanism, Transformers, BERT, and GPT: Tutorial and Survey 17 Nov 2020 · 0 repositories
-
DebateSum: A large-scale argument mining and summarization dataset 14 Nov 2020 · 3 repositories · arXiv:2011.07251Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Adapting a Language Model for Controlled Affective Text Generation 8 Nov 2020 · 1 repository · arXiv:2011.04000Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Semi-Supervised Low-Resource Style Transfer of Indonesian Informal to Formal Language with Iterative Forward-Translation 6 Nov 2020 · 1 repository · arXiv:2011.03286
-
Improving RNN transducer with normalized jointer network 3 Nov 2020 · 0 repositories · arXiv:2011.01576
-
Tabular Transformers for Modeling Multivariate Time Series 3 Nov 2020 · 1 repository · arXiv:2011.01843Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
The Amazing World of Neural Language Generation 1 Nov 2020 · 0 repositories
-
Visually-Grounded Planning without Vision: Language Models Infer Detailed Plans from High-level Instructions 1 Nov 2020 · 1 repository
-
Topic-Preserving Synthetic News Generation: An Adversarial Deep Reinforcement Learning Approach 30 Oct 2020 · 0 repositories · arXiv:2010.16324
-
Unsupervised Paraphrasing with Pretrained Language Models 24 Oct 2020 · 0 repositories · arXiv:2010.12885
-
LightSeq: A High Performance Inference Library for Transformers 23 Oct 2020 · 1 repository · arXiv:2010.13887
-
Topic Modeling with Contextualized Word Representation Clusters 23 Oct 2020 · 0 repositories · arXiv:2010.12626
-
Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset 22 Oct 2020 · 0 repositories · arXiv:2010.11395
-
Transferable Graph Optimizers for ML Compilers 21 Oct 2020 · 0 repositories · arXiv:2010.12438
-
Performance of Transfer Learning Model vs. Traditional Neural Network in Low System Resource Environment 20 Oct 2020 · 0 repositories · arXiv:2011.07962
-
Better Distractions: Transformer-based Distractor Generation and Multiple Choice Question Filtering 19 Oct 2020 · 0 repositories · arXiv:2010.09598
-
DA-Transformer: Distance-aware Transformer 14 Oct 2020 · 0 repositories · arXiv:2010.06925
-
Decoding Methods for Neural Narrative Generation 14 Oct 2020 · 2 repositories · arXiv:2010.07375