Methods › General › Regularization › Weight Decay › Papers, page 86
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 86 of 108: papers 8,501 to 8,600 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Explicit regularization and implicit bias in deep network classifiers trained with the square loss 31 Dec 2020 · 0 repositories · arXiv:2101.00072
-
KART: Parameterization of Privacy Leakage Scenarios from Pre-trained Language Models 31 Dec 2020 · 1 repository · arXiv:2101.00036
-
Fast WordPiece Tokenization 31 Dec 2020 · 1 repository · arXiv:2012.15524
-
Making Pre-trained Language Models Better Few-shot Learners 31 Dec 2020 · 9 repositories · arXiv:2012.15723Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 9 harvested samples) · 7 pointer-only (licence)
-
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers 31 Dec 2020 · 2 repositories · arXiv:2012.15828
-
The Pile: An 800GB Dataset of Diverse Text for Language Modeling 31 Dec 2020 · 22 repositories · arXiv:2101.00027Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Unified Mandarin TTS Front-end Based on Distilled BERT Model 31 Dec 2020 · 1 repository · arXiv:2012.15404
-
UNKs Everywhere: Adapting Multilingual Language Models to New Scripts 31 Dec 2020 · 2 repositories · arXiv:2012.15562
-
ECONET: Effective Continual Pretraining of Language Models for Event Temporal Reasoning 30 Dec 2020 · 2 repositories · arXiv:2012.15283
-
Deriving Contextualised Semantic Features from BERT (and Other Transformer Model) Embeddings 30 Dec 2020 · 0 repositories · arXiv:2012.15353
-
Improving BERT with Syntax-aware Local Attention 30 Dec 2020 · 1 repository · arXiv:2012.15150
-
Optimizing Deeper Transformers on Small Datasets 30 Dec 2020 · 1 repository · arXiv:2012.15355
-
Out of Order: How Important Is The Sequential Order of Words in a Sentence in Natural Language Understanding Tasks? 30 Dec 2020 · 0 repositories · arXiv:2012.15180
-
SemGloVe: Semantic Co-occurrences for GloVe from BERT 30 Dec 2020 · 3 repositories · arXiv:2012.15197
-
UnNatural Language Inference 30 Dec 2020 · 1 repository · arXiv:2101.00010
-
Reinforcement Learning for Control of Valves 29 Dec 2020 · 2 repositories · arXiv:2012.14668
-
Robust Dialogue Utterance Rewriting as Sequence Tagging 29 Dec 2020 · 1 repository · arXiv:2012.14535
-
A Paragraph-level Multi-task Learning Model for Scientific Fact-Verification 28 Dec 2020 · 1 repository · arXiv:2012.14500
-
BURT: BERT-inspired Universal Representation from Learning Meaningful Segment 28 Dec 2020 · 0 repositories · arXiv:2012.14320
-
Syntax-Enhanced Pre-trained Model 28 Dec 2020 · 1 repository · arXiv:2012.14116
-
A multi-task learning network using shared BERT models for aspect-based sentiment analysis 27 Dec 2020 · 0 repositories
-
ALP-KD: Attention-Based Layer Projection for Knowledge Distillation 27 Dec 2020 · 0 repositories · arXiv:2012.14022
-
An Embarrassingly Simple Model for Dialogue Relation Extraction 27 Dec 2020 · 1 repository · arXiv:2012.13873
-
Inserting Information Bottlenecks for Attribution in Transformers 27 Dec 2020 · 1 repository · arXiv:2012.13838
-
MeDAL: Medical Abbreviation Disambiguation Dataset for Natural Language Understanding Pretraining 27 Dec 2020 · 1 repository · arXiv:2012.13978
-
On the Granularity of Explanations in Model Agnostic NLP Interpretability 24 Dec 2020 · 1 repository · arXiv:2012.13189
-
To what extent do human explanations of model behavior align with actual model behavior? 24 Dec 2020 · 0 repositories · arXiv:2012.13354
-
Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing 23 Dec 2020 · 2 repositories · arXiv:2012.12627
-
Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge 23 Dec 2020 · 1 repository · arXiv:2012.12975
-
Training data-efficient image transformers & distillation through attention 23 Dec 2020 · 40 repositories · arXiv:2012.12877Syntology community repositories only · 12 ran (of which 1 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 19 harvested samples) · 3 pointer-only (licence)
-
Applying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages 22 Dec 2020 · 0 repositories · arXiv:2012.12121
-
Improved Biomedical Word Embeddings in the Transformer Era 22 Dec 2020 · 1 repository · arXiv:2012.11808
-
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning 22 Dec 2020 · 2 repositories · arXiv:2012.13255Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Learning to Retrieve Entity-Aware Knowledge and Generate Responses with Copy Mechanism for Task-Oriented Dialogue Systems 22 Dec 2020 · 1 repository · arXiv:2012.11937
-
Recognizing Emotion Cause in Conversations 22 Dec 2020 · 1 repository · arXiv:2012.11820
-
Uncertainty and Surprisal Jointly Deliver the Punchline: Exploiting Incongruity-Based Features for Humor Recognition 22 Dec 2020 · 0 repositories · arXiv:2012.12007
-
myGym: Modular Toolkit for Visuomotor Robotic Tasks 21 Dec 2020 · 0 repositories · arXiv:2012.11643
-
Cross-domain Retrieval in the Legal and Patent Domains: a Reproducibility Study 21 Dec 2020 · 1 repository · arXiv:2012.11405
-
Domain specific BERT representation for Named Entity Recognition of lab protocol 21 Dec 2020 · 1 repository · arXiv:2012.11145
-
Leveraging ParsBERT and Pretrained mT5 for Persian Abstractive Text Summarization 21 Dec 2020 · 1 repository · arXiv:2012.11204
-
Towards Incorporating Entity-specific Knowledge Graph Information in Predicting Drug-Drug Interactions 21 Dec 2020 · 0 repositories · arXiv:2012.11142
-
Breaking Writer's Block: Low-cost Fine-tuning of Natural Language Generation Models 19 Dec 2020 · 0 repositories · arXiv:2101.03216
-
An Empirical Study of Using Pre-trained BERT Models for Vietnamese Relation Extraction Task at VLSP 2020 18 Dec 2020 · 1 repository · arXiv:2012.10275
-
HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection 18 Dec 2020 · 6 repositories · arXiv:2012.10289
-
On Modality Bias in the TVQA Dataset 18 Dec 2020 · 1 repository · arXiv:2012.10210
-
A White Box Analysis of ColBERT 17 Dec 2020 · 0 repositories · arXiv:2012.09650
-
BERT Goes Shopping: Comparing Distributional Models for Product Representations 17 Dec 2020 · 1 repository · arXiv:2012.09807
-
Literature Retrieval for Precision Medicine with Neural Matching and Faceted Summarization 17 Dec 2020 · 1 repository · arXiv:2012.09355
-
MASKER: Masked Keyword Regularization for Reliable Text Classification 17 Dec 2020 · 1 repository · arXiv:2012.09392
-
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering 17 Dec 2020 · 0 repositories · arXiv:2012.09766
-
SceneFormer: Indoor Scene Generation with Transformers 17 Dec 2020 · 2 repositories · arXiv:2012.09793
-
A Lightweight Neural Model for Biomedical Entity Linking 16 Dec 2020 · 1 repository · arXiv:2012.08844
-
Learning from Mistakes: Using Mis-predictions as Harm Alerts in Language Pre-Training 16 Dec 2020 · 0 repositories · arXiv:2012.08789
-
Query expansion with artificially generated texts 16 Dec 2020 · 0 repositories · arXiv:2012.08787
-
R²-Net: Relation of Relation Learning Network for Sentence Semantic Matching 16 Dec 2020 · 0 repositories · arXiv:2012.08920
-
Revisiting Linformer with a modified self-attention with linear complexity 16 Dec 2020 · 0 repositories · arXiv:2101.10277
-
Pre-Training Transformers as Energy-Based Cloze Models 15 Dec 2020 · 1 repository · arXiv:2012.08561
-
RecipeNLG: A Cooking Recipes Dataset for Semi-Structured Text Generation 15 Dec 2020 · 1 repository
-
Extracting Training Data from Large Language Models 14 Dec 2020 · 3 repositories · arXiv:2012.07805Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding 14 Dec 2020 · 0 repositories · arXiv:2012.07335
-
Policy Gradient for items Recommendation on Virtual Taobao 14 Dec 2020 · 0 repositories
-
Policy Gradient RL Algorithms as Directed Acyclic Graphs 14 Dec 2020 · 1 repository · arXiv:2012.07763
-
Vartani Spellcheck -- Automatic Context-Sensitive Spelling Correction of OCR-generated Hindi Text Using BERT and Levenshtein Distance 14 Dec 2020 · 0 repositories · arXiv:2012.07652
-
Virtual Autonomous Driving with Reinforcement Learning 14 Dec 2020 · 0 repositories
-
Discriminative Pre-training for Low Resource Title Compression in Conversational Grocery 13 Dec 2020 · 0 repositories · arXiv:2012.06943
-
KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning 13 Dec 2020 · 0 repositories · arXiv:2012.07000
-
MiniVLM: A Smaller and Faster Vision-Language Model 13 Dec 2020 · 0 repositories · arXiv:2012.06946
-
CogALex-VI Shared Task: Transrelation - A Robust Multilingual Language Model for Multilingual Relation Identification 12 Dec 2020 · 1 repository
-
Yelp Review Rating Prediction: Machine Learning and Deep Learning Models 12 Dec 2020 · 1 repository · arXiv:2012.06690
-
Avoiding The Double Descent Phenomenon of Random Feature Models Using Hybrid Regularization 11 Dec 2020 · 1 repository · arXiv:2012.06667
-
Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment 11 Dec 2020 · 0 repositories · arXiv:2012.06373
-
Improving Task-Agnostic BERT Distillation with Layer Mapping Search 11 Dec 2020 · 0 repositories · arXiv:2012.06153
-
A Practical Approach towards Causality Mining in Clinical Text using Active Transfer Learning 10 Dec 2020 · 0 repositories · arXiv:2012.07563
-
As Good as New. How to Successfully Recycle English GPT-2 to Make Models for Other Languages 10 Dec 2020 · 1 repository · arXiv:2012.05628
-
Towards Neural Programming Interfaces 10 Dec 2020 · 1 repository · arXiv:2012.05983
-
Convex Regularization Behind Neural Reconstruction 9 Dec 2020 · 0 repositories · arXiv:2012.05169
-
Cross-lingual Word Sense Disambiguation using mBERT Embeddings with Syntactic Dependencies 9 Dec 2020 · 0 repositories · arXiv:2012.05300
-
Mapping the Space of Chemical Reactions Using Attention-Based Neural Networks 9 Dec 2020 · 1 repository · arXiv:2012.06051
-
Discourse Parsing of Contentious, Non-Convergent Online Discussions 8 Dec 2020 · 0 repositories · arXiv:2012.04585
-
Large-scale Quantitative Evidence of Media Impact on Public Opinion toward China 8 Dec 2020 · 0 repositories · arXiv:2012.07575
-
Efficient Estimation of Influence of a Training Instance 8 Dec 2020 · 0 repositories · arXiv:2012.04207
-
From Bag of Sentences to Document: Distantly Supervised Relation Extraction via Machine Reading Comprehension 8 Dec 2020 · 1 repository · arXiv:2012.04334
-
TADO: Time-varying Attention with Dual-Optimizer Model 8 Dec 2020 · 1 repository · arXiv:2012.04558
-
An Empirical Survey of Unsupervised Text Representation Methods on Twitter Data 7 Dec 2020 · 0 repositories · arXiv:2012.03468
-
CX DB8: A queryable extractive summarizer and semantic search engine 7 Dec 2020 · 2 repositories · arXiv:2012.03942
-
Dartmouth CS at WNUT-2020 Task 2: Informative COVID-19 Tweet Classification Using BERT 7 Dec 2020 · 0 repositories · arXiv:2012.04539
-
Detecting Insincere Questions from Text: A Transfer Learning Approach 7 Dec 2020 · 1 repository · arXiv:2012.07587
-
Efficient Reservoir Management through Deep Reinforcement Learning 7 Dec 2020 · 0 repositories · arXiv:2012.03822
-
KgPLM: Knowledge-guided Language Model Pre-training via Generative and Discriminative Learning 7 Dec 2020 · 0 repositories · arXiv:2012.03551
-
The Role of Regularization in Shaping Weight and Node Pruning Dependency and Dynamics 7 Dec 2020 · 0 repositories · arXiv:2012.03827
-
UBAR: Towards Fully End-to-End Task-Oriented Dialog Systems with GPT-2 7 Dec 2020 · 1 repository · arXiv:2012.03539
-
Enhanced Offensive Language Detection Through Data Augmentation 5 Dec 2020 · 0 repositories · arXiv:2012.02954
-
Pre-training Protein Language Models with Label-Agnostic Binding Pairs Enhances Performance in Downstream Tasks 5 Dec 2020 · 1 repository · arXiv:2012.03084
-
Automated Detection of Cyberbullying Against Women and Immigrants and Cross-domain Adaptability 4 Dec 2020 · 0 repositories · arXiv:2012.02565
-
EchoBERT: A Transformer-Based Approach for Behavior Detection in Echograms 4 Dec 2020 · 1 repository
-
Fine-tuning BERT for Low-Resource Natural Language Understanding via Active Learning 4 Dec 2020 · 0 repositories · arXiv:2012.02462
-
Modelling General Properties of Nouns by Selectively Averaging Contextualised Embeddings 4 Dec 2020 · 0 repositories · arXiv:2012.07580
-
Playing Text-Based Games with Common Sense 4 Dec 2020 · 0 repositories · arXiv:2012.02757
-
Pre-trained language models as knowledge bases for Automotive Complaint Analysis 4 Dec 2020 · 0 repositories · arXiv:2012.02558
-
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation 4 Dec 2020 · 0 repositories · arXiv:2012.02469