Methods › General › Learning Rate Schedules › Linear Warmup With Linear Decay › Papers, page 53
Linear Warmup With Linear Decay
Papers archive 2025-07-28
archive papers tagged: 7,076 · with a code link: 2,913 · where Syntology ran a sample: 650 (531 with a run with no instrument failure, 119 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,076 tagged: 531 with a run with no instrument failure, 119 where every run was a failure of Syntology's instrument)
Page 53 of 71: papers 5,201 to 5,300 of 7,076, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
BERT & Family Eat Word Salad: Experiments with Text Understanding 10 Jan 2021 · 1 repository · arXiv:2101.03453
-
Cisco at AAAI-CAD21 shared task: Predicting Emphasis in Presentation Slides using Contextualized Embeddings 10 Jan 2021 · 1 repository · arXiv:2101.11422
-
Learning Better Sentence Representation with Syntax Information 9 Jan 2021 · 0 repositories · arXiv:2101.03343
-
Contextual Non-Local Alignment over Full-Scale Representation for Text-Based Person Search 8 Jan 2021 · 2 repositories · arXiv:2101.03036
-
Misspelling Correction with Pre-trained Contextual Language Model 8 Jan 2021 · 0 repositories · arXiv:2101.03204
-
Applying Transfer Learning for Improving Domain-Specific Search Experience Using Query to Question Similarity 7 Jan 2021 · 0 repositories · arXiv:2101.02351
-
Exploring Text-transformers in AAAI 2021 Shared Task: COVID-19 Fake News Detection in English 7 Jan 2021 · 1 repository · arXiv:2101.02359
-
Homonym Identification using BERT -- Using a Clustering Approach 7 Jan 2021 · 0 repositories · arXiv:2101.02398
-
Transformer-based approach towards music emotion recognition from lyrics 6 Jan 2021 · 1 repository · arXiv:2101.02051
-
COVID-19: Comparative Analysis of Methods for Identifying Articles Related to Therapeutics and Vaccines without Using Labeled Data 5 Jan 2021 · 0 repositories · arXiv:2101.02017
-
I-BERT: Integer-only BERT Quantization 5 Jan 2021 · 7 repositories · arXiv:2101.01321
-
Improving reference mining in patents with BERT 4 Jan 2021 · 1 repository · arXiv:2101.01039
-
A Robust and Domain-Adaptive Approach for Low-Resource Named Entity Recognition 2 Jan 2021 · 1 repository · arXiv:2101.00388
-
CDLM: Cross-Document Language Modeling 2 Jan 2021 · 2 repositories · arXiv:2101.00406
-
End-to-End Training of Neural Retrievers for Open-Domain Question Answering 2 Jan 2021 · 2 repositories · arXiv:2101.00408
-
Lex-BERT: Enhancing BERT based NER with lexicons 2 Jan 2021 · 0 repositories · arXiv:2101.00396
-
Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex Words 2 Jan 2021 · 1 repository · arXiv:2101.00403Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 2 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
What all do audio transformer models hear? Probing Acoustic Representations for Language Delivery and its Structure 2 Jan 2021 · 0 repositories · arXiv:2101.00387
-
BROS: A Pre-trained Language Model for Understanding Texts in Document 1 Jan 2021 · 0 repositories
-
Cluster & Tune: Enhance BERT Performance in Low Resource Text Classification 1 Jan 2021 · 0 repositories
-
Cross-Probe BERT for Efficient and Effective Cross-Modal Search 1 Jan 2021 · 0 repositories
-
DACT-BERT: Increasing the efficiency and interpretability of BERT by using adaptive computation time. 1 Jan 2021 · 0 repositories
-
Data-aware Low-Rank Compression for Large NLP Models 1 Jan 2021 · 0 repositories
-
Deep Learning Proteins using a Triplet-BERT network 1 Jan 2021 · 0 repositories
-
Domain-slot Relationship Modeling using a Pre-trained Language Encoder for Multi-Domain Dialogue State Tracking 1 Jan 2021 · 0 repositories
-
Erasure for Advancing: Dynamic Self-Supervised Learning for Commonsense Reasoning 1 Jan 2021 · 0 repositories
-
EXPLORING VULNERABILITIES OF BERT-BASED APIS 1 Jan 2021 · 0 repositories
-
Isotropy in the Contextual Embedding Space: Clusters and Manifolds 1 Jan 2021 · 0 repositories
-
Modelling Drug-Target Binding Affinity using a BERT based Graph Neural network 1 Jan 2021 · 0 repositories
-
MULTI-SPAN QUESTION ANSWERING USING SPAN-IMAGE NETWORK 1 Jan 2021 · 0 repositories
-
On Explaining Your Explanations of BERT: An Empirical Study with Sequence Classification 1 Jan 2021 · 2 repositories · arXiv:2101.00196
-
Post-Training Weighted Quantization of Neural Networks for Language Models 1 Jan 2021 · 0 repositories
-
SkillBERT: “Skilling” the BERT to classify skills! 1 Jan 2021 · 0 repositories
-
Speeding up Deep Learning Training by Sharing Weights and Then Unsharing 1 Jan 2021 · 0 repositories
-
Syntactic Relevance XLNet Word Embedding Generation in Low-Resource Machine Translation 1 Jan 2021 · 0 repositories
-
Taking Notes on the Fly Helps Language Pre-Training 1 Jan 2021 · 0 repositories
-
Task-Agnostic and Adaptive-Size BERT Compression 1 Jan 2021 · 0 repositories
-
Towards Practical Second Order Optimization for Deep Learning 1 Jan 2021 · 0 repositories
-
U-BERT: Pre-training User Representations for Improved Recommendation 1 Jan 2021 · 0 repositories
-
UserBERT: Self-supervised User Representation Learning 1 Jan 2021 · 0 repositories
-
A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots Matters 31 Dec 2020 · 0 repositories · arXiv:2012.15682
-
A Multi-modal Deep Learning Model for Video Thumbnail Selection 31 Dec 2020 · 0 repositories · arXiv:2101.00073
-
Better Robustness by More Coverage: Adversarial Training with Mixup Augmentation for Robust Fine-tuning 31 Dec 2020 · 1 repository · arXiv:2012.15699
-
BinaryBERT: Pushing the Limit of BERT Quantization 31 Dec 2020 · 1 repository · arXiv:2012.15701Syntology official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
CoCoLM: COmplex COmmonsense Enhanced Language Model with Discourse Relations 31 Dec 2020 · 1 repository · arXiv:2012.15643
-
EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets 31 Dec 2020 · 1 repository · arXiv:2101.00063Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
KART: Parameterization of Privacy Leakage Scenarios from Pre-trained Language Models 31 Dec 2020 · 1 repository · arXiv:2101.00036
-
Fast WordPiece Tokenization 31 Dec 2020 · 1 repository · arXiv:2012.15524
-
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers 31 Dec 2020 · 2 repositories · arXiv:2012.15828
-
Unified Mandarin TTS Front-end Based on Distilled BERT Model 31 Dec 2020 · 1 repository · arXiv:2012.15404
-
UNKs Everywhere: Adapting Multilingual Language Models to New Scripts 31 Dec 2020 · 2 repositories · arXiv:2012.15562
-
ECONET: Effective Continual Pretraining of Language Models for Event Temporal Reasoning 30 Dec 2020 · 2 repositories · arXiv:2012.15283
-
Deriving Contextualised Semantic Features from BERT (and Other Transformer Model) Embeddings 30 Dec 2020 · 0 repositories · arXiv:2012.15353
-
Improving BERT with Syntax-aware Local Attention 30 Dec 2020 · 1 repository · arXiv:2012.15150
-
Optimizing Deeper Transformers on Small Datasets 30 Dec 2020 · 1 repository · arXiv:2012.15355
-
Out of Order: How Important Is The Sequential Order of Words in a Sentence in Natural Language Understanding Tasks? 30 Dec 2020 · 0 repositories · arXiv:2012.15180
-
SemGloVe: Semantic Co-occurrences for GloVe from BERT 30 Dec 2020 · 3 repositories · arXiv:2012.15197
-
UnNatural Language Inference 30 Dec 2020 · 1 repository · arXiv:2101.00010
-
A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification 28 Dec 2020 · 1 repository · arXiv:2012.14500
-
BURT: BERT-inspired Universal Representation from Learning Meaningful Segment 28 Dec 2020 · 0 repositories · arXiv:2012.14320
-
Syntax-Enhanced Pre-trained Model 28 Dec 2020 · 1 repository · arXiv:2012.14116
-
A multi-task learning network using shared BERT models for aspect-based sentiment analysis 27 Dec 2020 · 0 repositories
-
ALP-KD: Attention-Based Layer Projection for Knowledge Distillation 27 Dec 2020 · 0 repositories · arXiv:2012.14022
-
An Embarrassingly Simple Model for Dialogue Relation Extraction 27 Dec 2020 · 1 repository · arXiv:2012.13873
-
Inserting Information Bottlenecks for Attribution in Transformers 27 Dec 2020 · 1 repository · arXiv:2012.13838
-
MeDAL: Medical Abbreviation Disambiguation Dataset for Natural Language Understanding Pretraining 27 Dec 2020 · 1 repository · arXiv:2012.13978
-
On the Granularity of Explanations in Model Agnostic NLP Interpretability 24 Dec 2020 · 1 repository · arXiv:2012.13189
-
To what extent do human explanations of model behavior align with actual model behavior? 24 Dec 2020 · 0 repositories · arXiv:2012.13354
-
Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing 23 Dec 2020 · 2 repositories · arXiv:2012.12627
-
Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge 23 Dec 2020 · 1 repository · arXiv:2012.12975
-
Applying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages 22 Dec 2020 · 0 repositories · arXiv:2012.12121
-
Improved Biomedical Word Embeddings in the Transformer Era 22 Dec 2020 · 1 repository · arXiv:2012.11808
-
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning 22 Dec 2020 · 2 repositories · arXiv:2012.13255Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Learning to Retrieve Entity-Aware Knowledge and Generate Responses with Copy Mechanism for Task-Oriented Dialogue Systems 22 Dec 2020 · 1 repository · arXiv:2012.11937
-
Recognizing Emotion Cause in Conversations 22 Dec 2020 · 1 repository · arXiv:2012.11820
-
Cross-domain Retrieval in the Legal and Patent Domains: a Reproducibility Study 21 Dec 2020 · 1 repository · arXiv:2012.11405
-
Domain specific BERT representation for Named Entity Recognition of lab protocol 21 Dec 2020 · 1 repository · arXiv:2012.11145
-
Leveraging ParsBERT and Pretrained mT5 for Persian Abstractive Text Summarization 21 Dec 2020 · 1 repository · arXiv:2012.11204
-
Towards Incorporating Entity-specific Knowledge Graph Information in Predicting Drug-Drug Interactions 21 Dec 2020 · 0 repositories · arXiv:2012.11142
-
An Empirical Study of Using Pre-trained BERT Models for Vietnamese Relation Extraction Task at VLSP 2020 18 Dec 2020 · 1 repository · arXiv:2012.10275
-
HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection 18 Dec 2020 · 6 repositories · arXiv:2012.10289
-
On Modality Bias in the TVQA Dataset 18 Dec 2020 · 1 repository · arXiv:2012.10210
-
A White Box Analysis of ColBERT 17 Dec 2020 · 0 repositories · arXiv:2012.09650
-
BERT Goes Shopping: Comparing Distributional Models for Product Representations 17 Dec 2020 · 1 repository · arXiv:2012.09807
-
Literature Retrieval for Precision Medicine with Neural Matching and Faceted Summarization 17 Dec 2020 · 1 repository · arXiv:2012.09355
-
MASKER: Masked Keyword Regularization for Reliable Text Classification 17 Dec 2020 · 1 repository · arXiv:2012.09392
-
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering 17 Dec 2020 · 0 repositories · arXiv:2012.09766
-
SceneFormer: Indoor Scene Generation with Transformers 17 Dec 2020 · 2 repositories · arXiv:2012.09793
-
A Lightweight Neural Model for Biomedical Entity Linking 16 Dec 2020 · 1 repository · arXiv:2012.08844
-
DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion Recognition 16 Dec 2020 · 4 repositories · arXiv:2012.08695
-
Learning from Mistakes: Using Mis-predictions as Harm Alerts in Language Pre-Training 16 Dec 2020 · 0 repositories · arXiv:2012.08789
-
R²-Net: Relation of Relation Learning Network for Sentence Semantic Matching 16 Dec 2020 · 0 repositories · arXiv:2012.08920
-
Revisiting Linformer with a modified self-attention with linear complexity 16 Dec 2020 · 0 repositories · arXiv:2101.10277
-
Pre-Training Transformers as Energy-Based Cloze Models 15 Dec 2020 · 1 repository · arXiv:2012.08561
-
LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding 14 Dec 2020 · 0 repositories · arXiv:2012.07335
-
Vartani Spellcheck -- Automatic Context-Sensitive Spelling Correction of OCR-generated Hindi Text Using BERT and Levenshtein Distance 14 Dec 2020 · 0 repositories · arXiv:2012.07652
-
Discriminative Pre-training for Low Resource Title Compression in Conversational Grocery 13 Dec 2020 · 0 repositories · arXiv:2012.06943
-
KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning 13 Dec 2020 · 0 repositories · arXiv:2012.07000
-
MiniVLM: A Smaller and Faster Vision-Language Model 13 Dec 2020 · 0 repositories · arXiv:2012.06946
-
CogALex-VI Shared Task: Transrelation - A Robust Multilingual Language Model for Multilingual Relation Identification 12 Dec 2020 · 1 repository