Methods › General › Regularization › Weight Decay › Papers, page 61
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 61 of 108: papers 6,001 to 6,100 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Foundation Transformers 12 Oct 2022 · 4 repositories · arXiv:2210.06423Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
GMP*: Well-Tuned Gradual Magnitude Pruning Can Outperform Most BERT-Pruning Methods 12 Oct 2022 · 0 repositories · arXiv:2210.06384
-
On Text Style Transfer via Style Masked Language Models 12 Oct 2022 · 0 repositories · arXiv:2210.06394
-
Predictive Querying for Autoregressive Neural Sequence Models 12 Oct 2022 · 1 repository · arXiv:2210.06464Syntology official (archive's flag): 7 ran · 7 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples)
-
Probing Commonsense Knowledge in Pre-trained Language Models with Sense-level Precision and Expanded Vocabulary 12 Oct 2022 · 1 repository · arXiv:2210.06376
-
RankT5: Fine-Tuning T5 for Text Ranking with Ranking Losses 12 Oct 2022 · 0 repositories · arXiv:2210.10634
-
SUMBot: Summarizing Context in Open-Domain Dialogue Systems 12 Oct 2022 · 0 repositories · arXiv:2210.06496
-
A Win-win Deal: Towards Sparse and Robust Pre-trained Language Models 11 Oct 2022 · 1 repository · arXiv:2210.05211Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
An Exploration of Hierarchical Attention Transformers for Efficient Long Document Classification 11 Oct 2022 · 0 repositories · arXiv:2210.05529
-
CLIP also Understands Text: Prompting CLIP for Phrase Understanding 11 Oct 2022 · 0 repositories · arXiv:2210.05836
-
Multilingual BERT has an accent: Evaluating English influences on fluency in multilingual models 11 Oct 2022 · 0 repositories · arXiv:2210.05619
-
On the Interpolation of Contextualized Term-based Ranking with BM25 for Query-by-Example Retrieval 11 Oct 2022 · 1 repository · arXiv:2210.05512
-
On the Use of Semantically-Aligned Speech Representations for Spoken Language Understanding 11 Oct 2022 · 0 repositories · arXiv:2210.05291
-
Vote'n'Rank: Revision of Benchmarking with Social Choice Theory 11 Oct 2022 · 1 repository · arXiv:2210.05769Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 17 harvested samples)
-
DEPTWEET: A Typology for Social Media Texts to Detect Depression Severities 10 Oct 2022 · 1 repository · arXiv:2210.05372
-
Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter 10 Oct 2022 · 1 repository · arXiv:2210.04710
-
Multi-CLS BERT: An Efficient Alternative to Traditional Ensembling 10 Oct 2022 · 1 repository · arXiv:2210.05043
-
Multiagent Reinforcement Learning Based on Fusion-Multiactor-Attention-Critic for Multiple-Unmanned-Aerial-Vehicle Navigation Control 10 Oct 2022 · 1 repository
-
REV: Information-Theoretic Evaluation of Free-Text Rationales 10 Oct 2022 · 1 repository · arXiv:2210.04982Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
The Minimum Wage as an Anchor: Effects on Determinations of Fairness by Humans and AI 10 Oct 2022 · 0 repositories · arXiv:2210.10585
-
Uncertainty Quantification with Pre-trained Language Models: A Large-Scale Empirical Analysis 10 Oct 2022 · 1 repository · arXiv:2210.04714Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
ASDOT: Any-Shot Data-to-Text Generation with Pretrained Language Models 9 Oct 2022 · 1 repository · arXiv:2210.04325
-
Controllable Dialogue Simulation with In-Context Learning 9 Oct 2022 · 1 repository · arXiv:2210.04185
-
Fine-Grained Detection of Solidarity for Women and Migrants in 155 Years of German Parliamentary Debates 9 Oct 2022 · 2 repositories · arXiv:2210.04359
-
Fine-Tuning Pre-trained Transformers into Decaying Fast Weights 9 Oct 2022 · 1 repository · arXiv:2210.04243Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Better Pre-Training by Reducing Representation Confusion 9 Oct 2022 · 0 repositories · arXiv:2210.04246
-
Spread Love Not Hate: Undermining the Importance of Hateful Pre-training for Hate Speech Detection 9 Oct 2022 · 1 repository · arXiv:2210.04267
-
AlphaTuning: Quantization-Aware Parameter-Efficient Adaptation of Large-Scale Pre-Trained Language Models 8 Oct 2022 · 0 repositories · arXiv:2210.03858
-
KG-MTT-BERT: Knowledge Graph Enhanced BERT for Multi-Type Medical Text Classification 8 Oct 2022 · 0 repositories · arXiv:2210.03970
-
On Task-Adaptive Pretraining for Dialogue Response Selection 8 Oct 2022 · 0 repositories · arXiv:2210.04073
-
Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning 7 Oct 2022 · 0 repositories · arXiv:2210.03469
-
Automatic Chain of Thought Prompting in Large Language Models 7 Oct 2022 · 5 repositories · arXiv:2210.03493Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
DABERT: Dual Attention Enhanced BERT for Semantic Matching 7 Oct 2022 · 0 repositories · arXiv:2210.03454
-
Generating Quizzes to Support Training on Quality Management and Assurance in Space Science and Engineering 7 Oct 2022 · 0 repositories · arXiv:2210.03427
-
How Large Language Models are Transforming Machine-Paraphrased Plagiarism 7 Oct 2022 · 3 repositories · arXiv:2210.03568
-
Knowledge Injected Prompt Based Fine-tuning for Multi-label Few-shot ICD Coding 7 Oct 2022 · 1 repository · arXiv:2210.03304Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Measuring and Narrowing the Compositionality Gap in Language Models 7 Oct 2022 · 1 repository · arXiv:2210.03350
-
UU-Tax at SemEval-2022 Task 3: Improving the generalizability of language models for taxonomy classification through data augmentation 7 Oct 2022 · 1 repository · arXiv:2210.03378
-
PathProx: A Proximal Gradient Algorithm for Weight Decay Regularized Deep Neural Networks 6 Oct 2022 · 0 repositories · arXiv:2210.03069
-
Binding Language Models in Symbolic Languages 6 Oct 2022 · 4 repositories · arXiv:2210.02875Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
ByteTransformer: A High-Performance Transformer Boosted for Variable-Length Inputs 6 Oct 2022 · 1 repository · arXiv:2210.03052
-
Explainable Verbal Deception Detection using Transformers 6 Oct 2022 · 0 repositories · arXiv:2210.03080
-
Generalization Properties of Retrieval-based Models 6 Oct 2022 · 0 repositories · arXiv:2210.02617
-
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 6 Oct 2022 · 1 repository · arXiv:2210.02969
-
Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering 6 Oct 2022 · 1 repository · arXiv:2210.02627
-
Join-Chain Network: A Logical Reasoning View of the Multi-head Attention in Transformer 6 Oct 2022 · 0 repositories · arXiv:2210.02729
-
Matching Text and Audio Embeddings: Exploring Transfer-learning Strategies for Language-based Audio Retrieval 6 Oct 2022 · 0 repositories · arXiv:2210.02833
-
MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text 6 Oct 2022 · 0 repositories · arXiv:2210.02928
-
Rainier: Reinforced Knowledge Introspector for Commonsense Question Answering 6 Oct 2022 · 1 repository · arXiv:2210.03078Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 4 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples)
-
GLM-130B: An Open Bilingual Pre-trained Model 5 Oct 2022 · 9 repositories · arXiv:2210.02414Syntology official (archive's flag): 9 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 21 harvested samples)
-
Privacy-Preserving Text Classification on BERT Embeddings with Homomorphic Encryption 5 Oct 2022 · 0 repositories · arXiv:2210.02574
-
Revisiting Structured Dropout 5 Oct 2022 · 0 repositories · arXiv:2210.02570
-
Explaining Patterns in Data with Language Models via Interpretable Autoprompting 4 Oct 2022 · 2 repositories · arXiv:2210.01848Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
The (In)Effectiveness of Intermediate Task Training For Domain Adaptation and Cross-Lingual Transfer Learning 3 Oct 2022 · 0 repositories · arXiv:2210.01091
-
Complexity-Based Prompting for Multi-Step Reasoning 3 Oct 2022 · 0 repositories · arXiv:2210.00720
-
Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought 3 Oct 2022 · 2 repositories · arXiv:2210.01240Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Omnigrok: Grokking Beyond Algorithmic Data 3 Oct 2022 · 2 repositories · arXiv:2210.01117
-
Probing of Quantitative Values in Abstractive Summarization Models 3 Oct 2022 · 0 repositories · arXiv:2210.00667
-
Construction and Evaluation of a Self-Attention Model for Semantic Understanding of Sentence-Final Particles 1 Oct 2022 · 0 repositories · arXiv:2210.00282
-
LambdaKG: A Library for Pre-trained Language Model-Based Knowledge Graph Embeddings 1 Oct 2022 · 2 repositories · arXiv:2210.00305
-
Improving Robustness with Adaptive Weight Decay 30 Sep 2022 · 1 repository · arXiv:2210.00094Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun Resolution 30 Sep 2022 · 2 repositories · arXiv:2210.00131
-
Overparameterized ReLU Neural Networks Learn the Simplest Models: Neural Isometry and Exact Recovery 30 Sep 2022 · 1 repository · arXiv:2209.15265Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Scale-invariant Bayesian Neural Networks with Connectivity Tangent Kernel 30 Sep 2022 · 0 repositories · arXiv:2209.15208
-
SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation 30 Sep 2022 · 1 repository · arXiv:2209.15323
-
Bidirectional Language Models Are Also Few-shot Learners 29 Sep 2022 · 0 repositories · arXiv:2209.14500
-
Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning 29 Sep 2022 · 2 repositories · arXiv:2209.14610Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
NAG-GS: Semi-Implicit, Accelerated and Robust Stochastic Optimizer 29 Sep 2022 · 2 repositories · arXiv:2209.14937Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Neural Networks Efficiently Learn Low-Dimensional Representations with SGD 29 Sep 2022 · 0 repositories · arXiv:2209.14863
-
CEFER: A Four Facets Framework based on Context and Emotion embedded features for Implicit and Explicit Emotion Recognition 28 Sep 2022 · 0 repositories · arXiv:2209.13999
-
Downstream Datasets Make Surprisingly Good Pretraining Corpora 28 Sep 2022 · 1 repository · arXiv:2209.14389
-
Medical Image Captioning via Generative Pretrained Transformers 28 Sep 2022 · 0 repositories · arXiv:2209.13983
-
Supervised Contrastive Learning as Multi-Objective Optimization for Fine-Tuning Large Pre-trained Language Models 28 Sep 2022 · 0 repositories · arXiv:2209.14161
-
Who is GPT-3? An Exploration of Personality, Values and Demographics 28 Sep 2022 · 1 repository · arXiv:2209.14338
-
YATO: Yet Another deep learning based Text analysis Open toolkit 28 Sep 2022 · 1 repository · arXiv:2209.13877
-
How GPT-3 responds to different publics on climate change and Black Lives Matter: A critical appraisal of equity in conversational AI 27 Sep 2022 · 0 repositories · arXiv:2209.13627
-
Extractive Question Answering on Queries in Hindi and Tamil 27 Sep 2022 · 0 repositories · arXiv:2210.06356
-
Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models 27 Sep 2022 · 1 repository · arXiv:2209.13325Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
WikiDes: A Wikipedia-Based Dataset for Generating Short Descriptions from Paragraphs 27 Sep 2022 · 1 repository · arXiv:2209.13101
-
Do ever larger octopi still amplify reporting biases? Evidence from judgments of typical colour 26 Sep 2022 · 0 repositories · arXiv:2209.12786
-
News Summarization and Evaluation in the Era of GPT-3 26 Sep 2022 · 1 repository · arXiv:2209.12356
-
Towards Simple and Efficient Task-Adaptive Pre-training for Text Classification 26 Sep 2022 · 0 repositories · arXiv:2209.12943
-
SpeedLimit: Neural Architecture Search for Quantized Transformer Models 25 Sep 2022 · 0 repositories · arXiv:2209.12127
-
Sentiment Analysis on Inflation after Covid-19 25 Sep 2022 · 0 repositories · arXiv:2209.14737
-
Can Transformer Models Effectively Detect Software Aspects in StackOverflow Discussion? 24 Sep 2022 · 0 repositories · arXiv:2209.12065
-
Learning Chess With Language Models and Transformers 24 Sep 2022 · 0 repositories · arXiv:2209.11902
-
Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity 24 Sep 2022 · 0 repositories · arXiv:2209.12106
-
IDEA: Interactive DoublE Attentions from Label Embedding for Text Classification 23 Sep 2022 · 0 repositories · arXiv:2209.11407
-
A Case Report On The "A.I. Locked-In Problem": social concerns with modern NLP 22 Sep 2022 · 0 repositories · arXiv:2209.12687
-
Adaptation of domain-specific transformer models with text oversampling for sentiment analysis of social media posts on Covid-19 vaccines 22 Sep 2022 · 1 repository · arXiv:2209.10966
-
AIR-JPMC@SMM4H'22: Classifying Self-Reported Intimate Partner Violence in Tweets with Multiple BERT-based Models 22 Sep 2022 · 0 repositories · arXiv:2209.10763
-
DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation 22 Sep 2022 · 0 repositories · arXiv:2209.10797
-
Minimizing Human Assistance: Augmenting a Single Demonstration for Deep Reinforcement Learning 22 Sep 2022 · 0 repositories · arXiv:2209.11275
-
Bias at a Second Glance: A Deep Dive into Bias for German Educational Peer-Review Data Modeling 21 Sep 2022 · 2 repositories · arXiv:2209.10335Syntology official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
CAE: Mechanism to Diminish the Class Imbalanced in SLU Slot Filling Task 21 Sep 2022 · 1 repository
-
Representing Affect Information in Word Embeddings 21 Sep 2022 · 0 repositories · arXiv:2209.10583
-
Subject Verb Agreement Error Patterns in Meaningless Sentences: Humans vs. BERT 21 Sep 2022 · 0 repositories · arXiv:2209.10538
-
Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers 21 Sep 2022 · 0 repositories · arXiv:2209.10505
-
Towards Fine-tuning Pre-trained Language Models with Integer Forward and Backward Propagation 20 Sep 2022 · 0 repositories · arXiv:2209.09815
-
One-to-Many Semantic Communication Systems: Design, Implementation, Performance Evaluation 20 Sep 2022 · 0 repositories · arXiv:2209.09425