Methods › General › Regularization › Weight Decay › Papers, page 73
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 73 of 108: papers 7,201 to 7,300 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Pruning Attention Heads of Transformer Models Using A* Search: A Novel Approach to Compress Big NLP Architectures 28 Oct 2021 · 0 repositories · arXiv:2110.15225
-
Semi-Siamese Bi-encoder Neural Ranking Model Using Lightweight Fine-Tuning 28 Oct 2021 · 1 repository · arXiv:2110.14943
-
Anomaly-Injected Deep Support Vector Data Description for Text Outlier Detection 27 Oct 2021 · 0 repositories · arXiv:2110.14729
-
Automating Control of Overestimation Bias for Reinforcement Learning 26 Oct 2021 · 0 repositories · arXiv:2110.13523
-
CLAUSEREC: A Clause Recommendation Framework for AI-aided Contract Authoring 26 Oct 2021 · 0 repositories · arXiv:2110.15794
-
Post-processing for Individual Fairness 26 Oct 2021 · 1 repository · arXiv:2110.13796Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
s2s-ft: Fine-Tuning Pretrained Transformer Encoders for Sequence-to-Sequence Learning 26 Oct 2021 · 1 repository · arXiv:2110.13640
-
TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation 26 Oct 2021 · 1 repository · arXiv:2110.13412Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Fine-tuning of Pre-trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi 25 Oct 2021 · 1 repository · arXiv:2110.12687
-
Generating artificial texts as substitution or complement of training data 25 Oct 2021 · 0 repositories · arXiv:2110.13016
-
Recurrent Off-policy Baselines for Memory-based Continuous Control 25 Oct 2021 · 1 repository · arXiv:2110.12628
-
Paradigm Shift in Language Modeling: Revisiting CNN for Modeling Sanskrit Originated Bengali and Hindi Language 25 Oct 2021 · 0 repositories · arXiv:2110.13032
-
Hate and Offensive Speech Detection in Hindi and Marathi 23 Oct 2021 · 0 repositories · arXiv:2110.12200
-
Double Trouble: How to not explain a text classifier's decisions using counterfactuals synthesized by masked language models? 22 Oct 2021 · 1 repository · arXiv:2110.11929
-
Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering 22 Oct 2021 · 1 repository · arXiv:2110.11592
-
CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP 21 Oct 2021 · 1 repository · arXiv:2110.11316Syntology official (archive's flag): 4 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Fast Model Editing at Scale 21 Oct 2021 · 3 repositories · arXiv:2110.11309Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Modeling Performance in Open-Domain Dialogue with PARADISE 21 Oct 2021 · 0 repositories · arXiv:2110.11164
-
Computationally Efficient Safe Reinforcement Learning for Power Systems 20 Oct 2021 · 0 repositories · arXiv:2110.10333
-
Distributionally Robust Classifiers in Sentiment Analysis 20 Oct 2021 · 1 repository · arXiv:2110.10372
-
SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training 20 Oct 2021 · 0 repositories · arXiv:2110.10329
-
Ensemble ALBERT on SQuAD 2.0 19 Oct 2021 · 1 repository · arXiv:2110.09665
-
Risks of AI Foundation Models in Education 19 Oct 2021 · 0 repositories · arXiv:2110.10024
-
A Data Bootstrapping Recipe for Low Resource Multilingual Relation Classification 18 Oct 2021 · 0 repositories · arXiv:2110.09570
-
BERMo: What can BERT learn from ELMo? 18 Oct 2021 · 0 repositories · arXiv:2110.15802
-
Ceasing hate withMoH: Hate Speech Detection in Hindi-English Code-Switched Language 18 Oct 2021 · 0 repositories · arXiv:2110.09393
-
Contextual Hate Speech Detection in Code Mixed Text using Transformer Based Approaches 18 Oct 2021 · 0 repositories · arXiv:2110.09338
-
ViraPart: A Text Refinement Framework for Automatic Speech Recognition and Natural Language Processing Tasks in Persian 18 Oct 2021 · 0 repositories · arXiv:2110.09086
-
Illiterate DALL-E Learns to Compose 17 Oct 2021 · 1 repository · arXiv:2110.11405Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Reminding the Incremental Language Model via Data-Free Self-Distillation 17 Oct 2021 · 0 repositories · arXiv:2110.08745
-
Taming Visually Guided Sound Generation 17 Oct 2021 · 3 repositories · arXiv:2110.08791Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
A Short Study on Compressing Decoder-Based Language Models 16 Oct 2021 · 0 repositories · arXiv:2110.08460
-
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models 16 Oct 2021 · 0 repositories
-
EncT5: A Framework for Fine-tuning T5 as Non-autoregressive Models 16 Oct 2021 · 1 repository · arXiv:2110.08426
-
Evaluation of Transfer Learning for Polish with a text-to-text model 16 Oct 2021 · 0 repositories
-
Hierarchical Transformer Networks for Long-sequence and Multiple Clinical Documents Classification 16 Oct 2021 · 0 repositories
-
Hydra: A System for Large Multi-Model Deep Learning 16 Oct 2021 · 1 repository · arXiv:2110.08633
-
IMPLI: Investigating NLI Models' Performance on Figurative Language 16 Oct 2021 · 0 repositories
-
Knowledge Inheritance for Pre-trained Language Models 16 Oct 2021 · 0 repositories
-
Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens 16 Oct 2021 · 0 repositories
-
Old BERT, New Tricks: Artificial Language Learning for Pre-Trained Language Models 16 Oct 2021 · 1 repository
-
On the current state of reproducibility and reporting of uncertainty for Aspect-based Sentiment Analysis 16 Oct 2021 · 0 repositories
-
On the Robustness of Reading Comprehension Models to Entity Renaming 16 Oct 2021 · 1 repository · arXiv:2110.08555
-
PAGnol: An Extra-Large French Generative Model 16 Oct 2021 · 0 repositories · arXiv:2110.08554
-
PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization 16 Oct 2021 · 3 repositories · arXiv:2110.08499Syntology official (archive's flag): 4 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Semantic Search as Extractive Paraphrase Span Detection 16 Oct 2021 · 0 repositories
-
Sharpness-Aware Minimization Improves Language Model Generalization 16 Oct 2021 · 0 repositories · arXiv:2110.08529
-
SpelLM: Augmenting Chinese Spell Check Using Input Salience 16 Oct 2021 · 0 repositories
-
The Power of Prompt Tuning for Low-Resource Semantic Parsing 16 Oct 2021 · 0 repositories · arXiv:2110.08525
-
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models 16 Oct 2021 · 0 repositories
-
Robustness Challenges in Model Distillation and Pruning for Natural Language Understanding 16 Oct 2021 · 0 repositories · arXiv:2110.08419
-
Detecting Gender Bias in Transformer-based Models: A Case Study on BERT 15 Oct 2021 · 0 repositories · arXiv:2110.15733
-
Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining 15 Oct 2021 · 1 repository · arXiv:2110.08412Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Generating Natural Language Adversarial Examples through An Improved Beam Search Algorithm 15 Oct 2021 · 0 repositories · arXiv:2110.08036
-
Intent-based Product Collections for E-commerce using Pretrained Language Models 15 Oct 2021 · 0 repositories · arXiv:2110.08241
-
Kronecker Decomposition for GPT Compression 15 Oct 2021 · 0 repositories · arXiv:2110.08152
-
Probing as Quantifying Inductive Bias 15 Oct 2021 · 1 repository · arXiv:2110.08388
-
Tracing Origins: Coreference-aware Machine Reading Comprehension 15 Oct 2021 · 1 repository · arXiv:2110.07961
-
PARE: A Simple and Strong Baseline for Monolingual and Multilingual Distantly Supervised Relation Extraction 14 Oct 2021 · 1 repository · arXiv:2110.07415
-
bert2BERT: Towards Reusable Pretrained Language Models 14 Oct 2021 · 0 repositories · arXiv:2110.07143
-
BI-RADS BERT & Using Section Segmentation to Understand Radiology Reports 14 Oct 2021 · 1 repository · arXiv:2110.07552
-
Building Chinese Biomedical Language Models via Multi-Level Text Discrimination 14 Oct 2021 · 1 repository · arXiv:2110.07244
-
Interpreting the Robustness of Neural NLP Models to Textual Perturbations 14 Oct 2021 · 0 repositories · arXiv:2110.07159
-
Context-gloss Augmentation for Improving Word Sense Disambiguation 14 Oct 2021 · 0 repositories · arXiv:2110.07174
-
Can Machines Learn Morality? The Delphi Experiment 14 Oct 2021 · 1 repository · arXiv:2110.07574
-
Evaluating Off-the-Shelf Machine Listening and Natural Language Models for Automated Audio Captioning 14 Oct 2021 · 0 repositories · arXiv:2110.07410
-
Identifying Introductions in Podcast Episodes from Automatically Generated Transcripts 14 Oct 2021 · 1 repository · arXiv:2110.07096
-
P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts 14 Oct 2021 · 1 repository · arXiv:2110.07280
-
Symbolic Knowledge Distillation: from General Language Models to Commonsense Models 14 Oct 2021 · 1 repository · arXiv:2110.07178
-
Leveraging Generative Models for Covert Messaging: Challenges and Tradeoffs for "Dead-Drop" Deployments 13 Oct 2021 · 0 repositories · arXiv:2110.07009
-
Fake News Detection in Spanish Using Deep Learning Techniques 13 Oct 2021 · 1 repository · arXiv:2110.06461
-
Improving Graph-based Sentence Ordering with Iteratively Predicted Pairwise Orderings 13 Oct 2021 · 1 repository · arXiv:2110.06446Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Language Modelling via Learning to Rank 13 Oct 2021 · 0 repositories · arXiv:2110.06961
-
Maximizing Efficiency of Language Model Pre-training for Learning Representation 13 Oct 2021 · 0 repositories · arXiv:2110.06620
-
MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction 13 Oct 2021 · 1 repository · arXiv:2110.06651
-
Multistage linguistic conditioning of convolutional layers for speech emotion recognition 13 Oct 2021 · 0 repositories · arXiv:2110.06650
-
Parallel Deep Neural Networks Have Zero Duality Gap 13 Oct 2021 · 0 repositories · arXiv:2110.06482
-
Scaling Laws for the Few-Shot Adaptation of Pre-trained Image Classifiers 13 Oct 2021 · 0 repositories · arXiv:2110.06990
-
Towards Efficient NLP: A Standard Evaluation and A Strong Baseline 13 Oct 2021 · 1 repository · arXiv:2110.07038
-
Understanding of Emotion Perception from Art 13 Oct 2021 · 0 repositories · arXiv:2110.06486
-
Yformer: U-Net Inspired Transformer Architecture for Far Horizon Time Series Forecasting 13 Oct 2021 · 1 repository · arXiv:2110.08255
-
Regionalized models for Spanish language variations based on Twitter 12 Oct 2021 · 0 repositories · arXiv:2110.06128
-
ALL Dolphins Are Intelligent and SOME Are Friendly: Probing BERT for Nouns' Semantic Properties and their Prototypicality 12 Oct 2021 · 0 repositories · arXiv:2110.06376
-
BERTraffic: BERT-based Joint Speaker Role and Speaker Change Detection for Air Traffic Control Communications 12 Oct 2021 · 2 repositories · arXiv:2110.05781
-
Contrastive Learning for Representation Degeneration Problem in Sequential Recommendation 12 Oct 2021 · 2 repositories · arXiv:2110.05730Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Extracting Feelings of People Regarding COVID-19 by Social Network Mining 12 Oct 2021 · 0 repositories · arXiv:2110.06151
-
LightSeq2: Accelerated Training for Transformer-based Models on GPUs 12 Oct 2021 · 1 repository · arXiv:2110.05722
-
LiST: Lite Prompted Self-training Makes Parameter-Efficient Few-shot Learners 12 Oct 2021 · 1 repository · arXiv:2110.06274Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 2 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 4 pointer-only (licence)
-
Global Optimality Beyond Two Layers: Training Deep ReLU Networks via Convex Programs 11 Oct 2021 · 0 repositories · arXiv:2110.05518
-
Improving Gender Fairness of Pre-Trained Language Models without Catastrophic Forgetting 11 Oct 2021 · 0 repositories · arXiv:2110.05367
-
Multi-Task Learning for Situated Multi-Domain End-to-End Dialogue Systems 11 Oct 2021 · 0 repositories · arXiv:2110.05221
-
Offensive Language Detection with BERT-based models, By Customizing Attention Probabilities 11 Oct 2021 · 0 repositories · arXiv:2110.05133
-
Topic Modeling, Clade-assisted Sentiment Analysis, and Vaccine Brand Reputation Analysis of COVID-19 Vaccine-related Facebook Comments in the Philippines 11 Oct 2021 · 1 repository · arXiv:2111.04416
-
Towards Demystifying Representation Learning with Non-contrastive Self-supervision 11 Oct 2021 · 2 repositories · arXiv:2110.04947
-
PASTE: A Tagging-Free Decoding Framework Using Pointer Networks for Aspect Sentiment Triplet Extraction 10 Oct 2021 · 1 repository · arXiv:2110.04794
-
Yuan 1.0: Large-Scale Pre-trained Language Model in Zero-Shot and Few-Shot Learning 10 Oct 2021 · 1 repository · arXiv:2110.04725
-
An Isotropy Analysis in the Multilingual BERT Embedding Space 9 Oct 2021 · 1 repository · arXiv:2110.04504
-
Leveraging recent advances in Pre-Trained Language Models forEye-Tracking Prediction 9 Oct 2021 · 1 repository · arXiv:2110.04475
-
Pairwise Margin Maximization for Deep Neural Networks 9 Oct 2021 · 1 repository · arXiv:2110.04519
-
Vector-quantized Image Modeling with Improved VQGAN 9 Oct 2021 · 5 repositories · arXiv:2110.04627Syntology 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)