Methods › General › Learning Rate Schedules › Linear Warmup With Linear Decay › Papers, page 32
Linear Warmup With Linear Decay
Papers archive 2025-07-28
archive papers tagged: 7,076 · with a code link: 2,913 · where Syntology ran a sample: 650 (531 with a run with no instrument failure, 119 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,076 tagged: 531 with a run with no instrument failure, 119 where every run was a failure of Syntology's instrument)
Page 32 of 71: papers 3,101 to 3,200 of 7,076, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
PAC-MAN: Multi-Relation Network in Social Community for Personalized Hashtag Recommendation 14 Dec 2022 · 1 repository
-
DexBERT: Effective, Task-Agnostic and Fine-grained Representation Learning of Android Bytecode 12 Dec 2022 · 1 repository · arXiv:2212.05976
-
Classifying the Ideological Orientation of User-Submitted Texts in Social Media 12 Dec 2022 · 1 repository
-
Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin 10 Dec 2022 · 1 repository · arXiv:2212.05356
-
Incorporating Emotions into Health Mention Classification Task on Social Media 9 Dec 2022 · 1 repository · arXiv:2212.05039
-
Explain to me like I am five -- Sentence Simplification Using Transformers 8 Dec 2022 · 1 repository · arXiv:2212.04595
-
Memorization of Named Entities in Fine-tuned BERT Models 7 Dec 2022 · 1 repository · arXiv:2212.03749
-
Learning-To-Embed: Adopting Transformer based models for E-commerce Products Representation Learning 7 Dec 2022 · 0 repositories · arXiv:2212.03725
-
SimVTP: Simple Video Text Pre-training with Masked Autoencoders 7 Dec 2022 · 0 repositories · arXiv:2212.03490
-
TweetDrought: A Deep-Learning Drought Impacts Recognizer based on Twitter Data 7 Dec 2022 · 0 repositories · arXiv:2212.04001
-
CySecBERT: A Domain-Adapted Language Model for the Cybersecurity Domain 6 Dec 2022 · 0 repositories · arXiv:2212.02974
-
Vision Transformer Computation and Resilience for Dynamic Inference 6 Dec 2022 · 0 repositories · arXiv:2212.02687
-
LUNA: Language Understanding with Number Augmentations on Transformers via Number Plugins and Pre-training 6 Dec 2022 · 1 repository · arXiv:2212.02691Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Modern French Poetry Generation with RoBERTa and GPT-2 6 Dec 2022 · 0 repositories · arXiv:2212.02911
-
Style transfer and classification in hebrew news items 6 Dec 2022 · 0 repositories · arXiv:2212.03019
-
Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog 5 Dec 2022 · 0 repositories · arXiv:2212.02168
-
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning 2 Dec 2022 · 0 repositories · arXiv:2212.01378
-
Event knowledge in large language models: the gap between the impossible and the unlikely 2 Dec 2022 · 1 repository · arXiv:2212.01488
-
Adapted Multimodal BERT with Layer-wise Fusion for Sentiment Analysis 1 Dec 2022 · 0 repositories · arXiv:2212.00678
-
BudgetLongformer: Can we Cheaply Pretrain a SotA Legal Language Model From Scratch? 30 Nov 2022 · 0 repositories · arXiv:2211.17135
-
ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT 30 Nov 2022 · 1 repository · arXiv:2211.17201Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
HEAT: Hardware-Efficient Automatic Tensor Decomposition for Transformer Compression 30 Nov 2022 · 0 repositories · arXiv:2211.16749
-
Composition based oxidation state prediction of materials using deep learning 29 Nov 2022 · 1 repository · arXiv:2211.15895
-
Diverse Multi-Answer Retrieval with Determinantal Point Processes 29 Nov 2022 · 0 repositories · arXiv:2211.16029
-
Outfit Generation and Recommendation -- An Experimental Study 29 Nov 2022 · 0 repositories · arXiv:2211.16353
-
Automatically Extracting Information in Medical Dialogue: Expert System And Attention for Labelling 28 Nov 2022 · 0 repositories · arXiv:2211.15544
-
DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models 28 Nov 2022 · 1 repository · arXiv:2211.15029Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Handling and extracting key entities from customer conversations using Speech recognition and Named Entity recognition 28 Nov 2022 · 0 repositories · arXiv:2211.17107
-
Is it Required? Ranking the Skills Required for a Job-Title 28 Nov 2022 · 0 repositories · arXiv:2212.08553
-
Revisiting Distance Metric Learning for Few-Shot Natural Language Classification 28 Nov 2022 · 0 repositories · arXiv:2211.15202
-
Scientific and Creative Analogies in Pretrained Language Models 28 Nov 2022 · 2 repositories · arXiv:2211.15268
-
ESIE-BERT: Enriching Sub-words Information Explicitly with BERT for Joint Intent Classification and SlotFilling 27 Nov 2022 · 0 repositories · arXiv:2211.14829
-
Understanding BLOOM: An empirical study on diverse NLP tasks 27 Nov 2022 · 0 repositories · arXiv:2211.14865
-
An Analysis of Social Biases Present in BERT Variants Across Multiple Languages 25 Nov 2022 · 1 repository · arXiv:2211.14402Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Finetuning BERT on Partially Annotated NER Corpora 25 Nov 2022 · 1 repository · arXiv:2211.14360
-
InDEX: Indonesian Idiom and Expression Dataset for Cloze Test 24 Nov 2022 · 0 repositories · arXiv:2211.13376
-
Using Selective Masking as a Bridge between Pre-training and Fine-tuning 24 Nov 2022 · 0 repositories · arXiv:2211.13815
-
Holistic Visual-Textual Sentiment Analysis with Prior Models 23 Nov 2022 · 1 repository · arXiv:2211.12981
-
SEAT: Stable and Explainable Attention 23 Nov 2022 · 0 repositories · arXiv:2211.13290
-
Word-Level Representation From Bytes For Language Modeling 23 Nov 2022 · 0 repositories · arXiv:2211.12677
-
OLGA : An Ontology and LSTM-based approach for generating Arithmetic Word Problems (AWPs) of transfer type 22 Nov 2022 · 0 repositories · arXiv:2211.12164
-
AF Adapter: Continual Pretraining for Building Chinese Biomedical Language Model 21 Nov 2022 · 1 repository · arXiv:2211.11363
-
Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task 21 Nov 2022 · 2 repositories · arXiv:2211.11216
-
L3Cube-HindBERT and DevBERT: Pre-Trained BERT Transformer models for Devanagari based Hindi and Marathi Languages 21 Nov 2022 · 0 repositories · arXiv:2211.11418
-
L3Cube-MahaSBERT and HindSBERT: Sentence BERT Models and Benchmarking BERT Sentence Representations for Hindi and Marathi 21 Nov 2022 · 1 repository · arXiv:2211.11187
-
TCBERT: A Technical Report for Chinese Topic Classification BERT 21 Nov 2022 · 0 repositories · arXiv:2211.11304
-
Conceptor-Aided Debiasing of Large Language Models 20 Nov 2022 · 0 repositories · arXiv:2211.11087
-
Detecting Conspiracy Theory Against COVID-19 Vaccines 20 Nov 2022 · 0 repositories · arXiv:2211.13003
-
Feature Weaken: Vicinal Data Augmentation for Classification 20 Nov 2022 · 0 repositories · arXiv:2211.10944
-
Understanding and Improving Knowledge Distillation for Quantization-Aware Training of Large Transformer Encoders 20 Nov 2022 · 1 repository · arXiv:2211.11014
-
A survey on knowledge-enhanced multimodal learning 19 Nov 2022 · 0 repositories · arXiv:2211.12328
-
Entity-Assisted Language Models for Identifying Check-worthy Sentences 19 Nov 2022 · 0 repositories · arXiv:2211.10678
-
Leveraging Users' Social Network Embeddings for Fake News Detection on Twitter 19 Nov 2022 · 0 repositories · arXiv:2211.10672
-
Metadata Might Make Language Models Better 18 Nov 2022 · 0 repositories · arXiv:2211.10086
-
Where did you tweet from? Inferring the origin locations of tweets based on contextual information 18 Nov 2022 · 0 repositories · arXiv:2211.16506
-
LongFNT: Long-form Speech Recognition with Factorized Neural Transducer 17 Nov 2022 · 0 repositories · arXiv:2211.09412
-
ProtSi: Prototypical Siamese Network with Data Augmentation for Few-Shot Subjective Answer Evaluation 17 Nov 2022 · 1 repository · arXiv:2211.09855
-
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers 17 Nov 2022 · 1 repository · arXiv:2211.11586
-
Fast and Accurate FSA System Using ELBERT: An Efficient and Lightweight BERT 16 Nov 2022 · 0 repositories · arXiv:2211.08842
-
An FNet based Auto Encoder for Long Sequence News Story Generation 15 Nov 2022 · 1 repository · arXiv:2211.08295
-
Empowering Language Models with Knowledge Graph Reasoning for Question Answering 15 Nov 2022 · 0 repositories · arXiv:2211.08380
-
RobBERT-2022: Updating a Dutch Language Model to Account for Evolving Language Use 15 Nov 2022 · 0 repositories · arXiv:2211.08192
-
GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No Cost 13 Nov 2022 · 1 repository · arXiv:2211.06993
-
Xu at SemEval-2022 Task 4: Pre-BERT Neural Network Methods vs Post-BERT RoBERTa Approach for Patronizing and Condescending Language Detection 13 Nov 2022 · 1 repository · arXiv:2211.06874
-
Dark patterns in e-commerce: a dataset and its baseline evaluations 12 Nov 2022 · 1 repository · arXiv:2211.06543
-
Using Persuasive Writing Strategies to Explain and Detect Health Misinformation 11 Nov 2022 · 1 repository · arXiv:2211.05985
-
BERT-Based Combination of Convolutional and Recurrent Neural Network for Indonesian Sentiment Analysis 10 Nov 2022 · 0 repositories · arXiv:2211.05273
-
BERT in Plutarch's Shadows 10 Nov 2022 · 0 repositories · arXiv:2211.05673
-
Biomedical Multi-hop Question Answering Using Knowledge Graph Embeddings and Language Models 10 Nov 2022 · 0 repositories · arXiv:2211.05351
-
PAD-Net: An Efficient Framework for Dynamic Networks 10 Nov 2022 · 1 repository · arXiv:2211.05528
-
Syntax-Guided Domain Adaptation for Aspect-based Sentiment Analysis 10 Nov 2022 · 0 repositories · arXiv:2211.05457
-
Collateral facilitation in humans and language models 9 Nov 2022 · 1 repository · arXiv:2211.05198Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Cross-lingual Transfer Learning for Check-worthy Claim Identification over Twitter 9 Nov 2022 · 0 repositories · arXiv:2211.05087
-
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token 9 Nov 2022 · 1 repository · arXiv:2211.04898
-
Sentiment Analysis of Persian Language: Review of Algorithms, Approaches and Datasets 9 Nov 2022 · 0 repositories · arXiv:2212.06041
-
A Multimodal Approach for Dementia Detection from Spontaneous Speech with Tensor Fusion Layer 8 Nov 2022 · 0 repositories · arXiv:2211.04368
-
Discover, Explanation, Improvement: An Automatic Slice Detection Framework for Natural Language Processing 8 Nov 2022 · 0 repositories · arXiv:2211.04476
-
AD-BERT: Using Pre-trained contextualized embeddings to Predict the Progression from Mild Cognitive Impairment to Alzheimer's Disease 7 Nov 2022 · 0 repositories · arXiv:2212.06042
-
Suffix Retrieval-Augmented Language Modeling 6 Nov 2022 · 1 repository · arXiv:2211.03053
-
BERT-Deep CNN: State-of-the-Art for Sentiment Analysis of COVID-19 Tweets 4 Nov 2022 · 0 repositories · arXiv:2211.09733
-
BERT for Long Documents: A Case Study of Automated ICD Coding 4 Nov 2022 · 0 repositories · arXiv:2211.02519
-
Continuous Prompt Tuning Based Textual Entailment Model for E-commerce Entity Typing 4 Nov 2022 · 1 repository · arXiv:2211.02483
-
Fine-Tuning Language Models via Epistemic Neural Networks 3 Nov 2022 · 1 repository · arXiv:2211.01568Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
BECTRA: Transducer-based End-to-End ASR with BERT-Enhanced Encoder 2 Nov 2022 · 0 repositories · arXiv:2211.00792
-
Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language Model 2 Nov 2022 · 0 repositories · arXiv:2211.01200
-
Processing Long Legal Documents with Pre-trained Transformers: Modding LegalBERT and Longformer 2 Nov 2022 · 0 repositories · arXiv:2211.00974
-
ClassActionPrediction: A Challenging Benchmark for Legal Judgment Prediction of Class Action Cases in the US 1 Nov 2022 · 1 repository · arXiv:2211.00582
-
Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features 1 Nov 2022 · 0 repositories · arXiv:2211.00342
-
Reduce, Reuse, Recycle: Improving Training Efficiency with Distillation 1 Nov 2022 · 0 repositories · arXiv:2211.00683
-
Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization 31 Oct 2022 · 1 repository · arXiv:2210.17170
-
Leveraging Pre-trained Models for Failure Analysis Triplets Generation 31 Oct 2022 · 0 repositories · arXiv:2210.17497
-
QuaLA-MiniLM: a Quantized Length Adaptive MiniLM 31 Oct 2022 · 2 repositories · arXiv:2210.17114
-
SDCL: Self-Distillation Contrastive Learning for Chinese Spell Checking 31 Oct 2022 · 0 repositories · arXiv:2210.17168
-
Parameter-Efficient Tuning Makes a Good Classification Head 30 Oct 2022 · 1 repository · arXiv:2210.16771
-
BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model 29 Oct 2022 · 0 repositories · arXiv:2210.16663
-
Empirical Evaluation of Post-Training Quantization Methods for Language Tasks 29 Oct 2022 · 0 repositories · arXiv:2210.16621
-
Exploiting prompt learning with pre-trained language models for Alzheimer's Disease detection 29 Oct 2022 · 1 repository · arXiv:2210.16539
-
BEBERT: Efficient and Robust Binary Ensemble BERT 28 Oct 2022 · 1 repository · arXiv:2210.15976
-
Feature Engineering vs BERT on Twitter Data 28 Oct 2022 · 0 repositories · arXiv:2210.16168
-
On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment Analysis 28 Oct 2022 · 0 repositories · arXiv:2210.15937