Methods › General › Regularization › Weight Decay › Papers, page 57
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 57 of 108: papers 5,601 to 5,700 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
What do Language Models know about word senses? Zero-Shot WSD with Language Models and Domain Inventories 7 Feb 2023 · 0 repositories · arXiv:2302.03353
-
What Matters In The Structured Pruning of Generative Language Models? 7 Feb 2023 · 1 repository · arXiv:2302.03773Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples)
-
Context-Gloss Augmentation for Improving Arabic Target Sense Verification 6 Feb 2023 · 0 repositories · arXiv:2302.03126
-
cross-modal fusion techniques for utterance-level emotion recognition from text and speech 5 Feb 2023 · 0 repositories · arXiv:2302.02447
-
Nationality Bias in Text Generation 5 Feb 2023 · 0 repositories · arXiv:2302.02463
-
Quantized Distributed Training of Large Models with Convergence Guarantees 5 Feb 2023 · 0 repositories · arXiv:2302.02390
-
VuLASTE: Long Sequence Model with Abstract Syntax Tree Embedding for vulnerability Detection 5 Feb 2023 · 0 repositories · arXiv:2302.02345
-
A New cross-domain strategy based XAI models for fake news detection 4 Feb 2023 · 0 repositories · arXiv:2302.02122
-
Deep Reinforcement Learning for Traffic Light Control in Intelligent Transportation Systems 4 Feb 2023 · 0 repositories · arXiv:2302.03669
-
Knowledge Graph Completion Method Combined With Adaptive Enhanced Semantic Information 4 Feb 2023 · 0 repositories · arXiv:2302.02116
-
REaLTabFormer: Generating Realistic Relational and Tabular Data using Transformers 4 Feb 2023 · 3 repositories · arXiv:2302.02041Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Evaluating Large Language Models in Theory of Mind Tasks 4 Feb 2023 · 0 repositories · arXiv:2302.02083
-
ANTM: An Aligned Neural Topic Model for Exploring Evolving Topics 3 Feb 2023 · 1 repository · arXiv:2302.01501
-
Bioformer: an efficient transformer language model for biomedical text mining 3 Feb 2023 · 1 repository · arXiv:2302.01588
-
Detecting Reddit Users with Depression Using a Hybrid Neural Network SBERT-CNN 3 Feb 2023 · 0 repositories · arXiv:2302.02759
-
Revisiting Intermediate Layer Distillation for Compressing Language Models: An Overfitting Perspective 3 Feb 2023 · 1 repository · arXiv:2302.01530
-
Towards Few-Shot Identification of Morality Frames using In-Context Learning 3 Feb 2023 · 0 repositories · arXiv:2302.02029
-
Creating a Large Language Model of a Philosopher 2 Feb 2023 · 0 repositories · arXiv:2302.01339
-
Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment 2 Feb 2023 · 1 repository · arXiv:2302.00902
-
Longformer: Longitudinal Transformer for Alzheimer's Disease Classification with Structural MRIs 2 Feb 2023 · 1 repository · arXiv:2302.00901
-
Mnemosyne: Learning to Train Transformers with Transformers 2 Feb 2023 · 0 repositories · arXiv:2302.01128
-
Resilient Binary Neural Network 2 Feb 2023 · 1 repository · arXiv:2302.00956Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Semantic Coherence Markers for the Early Diagnosis of the Alzheimer Disease 2 Feb 2023 · 1 repository · arXiv:2302.01025
-
Large language models predict human sensory judgments across six modalities 2 Feb 2023 · 0 repositories · arXiv:2302.01308
-
An Empirical Study on the Transferability of Transformer Modules in Parameter-Efficient Fine-Tuning 1 Feb 2023 · 0 repositories · arXiv:2302.00378
-
Analyzing Leakage of Personally Identifiable Information in Language Models 1 Feb 2023 · 1 repository · arXiv:2302.00539
-
Co-Writing with Opinionated Language Models Affects Users' Views 1 Feb 2023 · 0 repositories · arXiv:2302.00560
-
Improving Few-Shot Generalization by Exploring and Exploiting Auxiliary Data 1 Feb 2023 · 1 repository · arXiv:2302.00674
-
Numeracy from Literacy: Data Science as an Emergent Skill from Large Language Models 31 Jan 2023 · 0 repositories · arXiv:2301.13382
-
The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments 31 Jan 2023 · 1 repository · arXiv:2301.13771
-
Adaptive Machine Translation with Large Language Models 30 Jan 2023 · 1 repository · arXiv:2301.13294
-
CSDR-BERT: a pre-trained scientific dataset match model for Chinese Scientific Dataset Retrieval 30 Jan 2023 · 0 repositories · arXiv:2301.12700
-
REPLUG: Retrieval-Augmented Black-Box Language Models 30 Jan 2023 · 3 repositories · arXiv:2301.12652Syntology 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Representation biases in sentence transformers 30 Jan 2023 · 0 repositories · arXiv:2301.13039
-
Specializing Smaller Language Models towards Multi-Step Reasoning 30 Jan 2023 · 2 repositories · arXiv:2301.12726Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
V2N Service Scaling with Deep Reinforcement Learning 30 Jan 2023 · 0 repositories · arXiv:2301.13324
-
A Discerning Several Thousand Judgments: GPT-3 Rates the Article + Adjective + Numeral + Noun Construction 29 Jan 2023 · 0 repositories · arXiv:2301.12564
-
BERT-based Authorship Attribution on the Romanian Dataset called ROST 29 Jan 2023 · 0 repositories · arXiv:2301.12500
-
Global Flood Prediction: a Multimodal Machine Learning Approach 29 Jan 2023 · 0 repositories · arXiv:2301.12548
-
Semantic Tagging with LSTM-CRF 28 Jan 2023 · 0 repositories · arXiv:2301.12206
-
Towards Equitable Representation in Text-to-Image Synthesis Models with the Cross-Cultural Understanding Benchmark (CCUB) Dataset 28 Jan 2023 · 1 repository · arXiv:2301.12073
-
A Comparative Study of Pretrained Language Models for Long Clinical Text 27 Jan 2023 · 1 repository · arXiv:2301.11847
-
A Multi-View Joint Learning Framework for Embedding Clinical Codes and Text Using Graph Neural Networks 27 Jan 2023 · 0 repositories · arXiv:2301.11608
-
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU? 27 Jan 2023 · 0 repositories · arXiv:2301.11688
-
Context Matters: A Strategy to Pre-train Language Model for Science Education 27 Jan 2023 · 0 repositories · arXiv:2301.12031
-
Predicting Sentence-Level Factuality of News and Bias of Media Outlets 27 Jan 2023 · 1 repository · arXiv:2301.11850
-
The Exploration of Knowledge-Preserving Prompts for Document Summarisation 27 Jan 2023 · 0 repositories · arXiv:2301.11719
-
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning 27 Jan 2023 · 1 repository · arXiv:2301.11916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
ThoughtSource: A central hub for large language model reasoning data 27 Jan 2023 · 1 repository · arXiv:2301.11596Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases 27 Jan 2023 · 1 repository · arXiv:2301.12017
-
Understanding the Effectiveness of Very Large Language Models on Dialog Evaluation 27 Jan 2023 · 0 repositories · arXiv:2301.12004
-
A benchmark for toxic comment classification on Civil Comments dataset 26 Jan 2023 · 1 repository · arXiv:2301.11125
-
BERT-Embedding and Citation Network Analysis based Query Expansion Technique for Scholarly Search 26 Jan 2023 · 0 repositories · arXiv:2301.11069
-
Causal Reasoning of Entities and Events in Procedural Texts 26 Jan 2023 · 1 repository · arXiv:2301.10896Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 15 harvested samples)
-
ExaRanker: Explanation-Augmented Neural Ranker 25 Jan 2023 · 1 repository · arXiv:2301.10521
-
A Novel Deep Reinforcement Learning-based Approach for Enhancing Spectral Efficiency of IRS-assisted Wireless Systems 24 Jan 2023 · 0 repositories · arXiv:2302.14706
-
A Stability Analysis of Fine-Tuning a Pre-Trained Model 24 Jan 2023 · 0 repositories · arXiv:2301.09820
-
Audience-Centric Natural Language Generation via Style Infusion 24 Jan 2023 · 1 repository · arXiv:2301.10283
-
The Next Chapter: A Study of Large Language Models in Storytelling 24 Jan 2023 · 0 repositories · arXiv:2301.09790
-
Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards 24 Jan 2023 · 0 repositories · arXiv:2301.10095
-
Large language models can segment narrative events similarly to humans 24 Jan 2023 · 0 repositories · arXiv:2301.10297
-
Multitask Instruction-based Prompting for Fallacy Recognition 24 Jan 2023 · 0 repositories · arXiv:2301.09992
-
AI model GPT-3 (dis)informs us better than humans 23 Jan 2023 · 0 repositories · arXiv:2301.11924
-
Deep Learning Meets Sparse Regularization: A Signal Processing Perspective 23 Jan 2023 · 0 repositories · arXiv:2301.09554
-
Injecting the BM25 Score as Text Improves BERT-Based Re-rankers 23 Jan 2023 · 1 repository · arXiv:2301.09728
-
StockEmotions: Discover Investor Emotions for Financial Sentiment Analysis and Multivariate Time Series 23 Jan 2023 · 2 repositories · arXiv:2301.09279
-
Stress Test for BERT and Deep Models: Predicting Words from Italian Poetry 21 Jan 2023 · 0 repositories · arXiv:2302.09303
-
SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction 21 Jan 2023 · 0 repositories · arXiv:2301.08984
-
Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine 20 Jan 2023 · 1 repository · arXiv:2301.08745
-
On Multi-Agent Deep Deterministic Policy Gradients and their Explainability for SMARTS Environment 20 Jan 2023 · 0 repositories · arXiv:2301.09420
-
Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions 20 Jan 2023 · 2 repositories · arXiv:2301.08810Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Which Features are Learned by CodeBert: An Empirical Study of the BERT-based Source Code Representation Learning 20 Jan 2023 · 0 repositories · arXiv:2301.08427
-
Batch Prompting: Efficient Inference with Large Language Model APIs 19 Jan 2023 · 2 repositories · arXiv:2301.08721
-
A Novel Sparse Regularizer 18 Jan 2023 · 0 repositories · arXiv:2301.07285
-
An Error-Guided Correction Model for Chinese Spelling Error Correction 16 Jan 2023 · 1 repository · arXiv:2301.06323
-
TEDB System Description to a Shared Task on Euphemism Detection 2022 16 Jan 2023 · 1 repository · arXiv:2301.06602
-
Improving Noise Robustness for Spoken Content Retrieval using Semi-supervised ASR and N-best Transcripts for BERT-based Ranking Models 15 Jan 2023 · 0 repositories · arXiv:2301.06056
-
T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations 15 Jan 2023 · 1 repository · arXiv:2301.06052Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
GPT as Knowledge Worker: A Zero-Shot Evaluation of (AI)CPA Capabilities 11 Jan 2023 · 1 repository · arXiv:2301.04408
-
NarrowBERT: Accelerating Masked Language Model Pretraining and Inference 11 Jan 2023 · 1 repository · arXiv:2301.04761Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Topics in Contextualised Attention Embeddings 11 Jan 2023 · 0 repositories · arXiv:2301.04339
-
Language Models sounds the Death Knell of Knowledge Graphs 10 Jan 2023 · 0 repositories · arXiv:2301.03980
-
Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language Models 10 Jan 2023 · 0 repositories · arXiv:2301.03797
-
Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling 9 Jan 2023 · 2 repositories · arXiv:2301.03580Syntology official (archive's flag): 9 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Online Fake Review Detection Using Supervised Machine Learning And BERT Model 9 Jan 2023 · 0 repositories · arXiv:2301.03225
-
Automatic Generation of German Drama Texts Using Fine Tuned GPT-2 Models 8 Jan 2023 · 0 repositories · arXiv:2301.03119
-
App Review Driven Collaborative Bug Finding 7 Jan 2023 · 1 repository · arXiv:2301.02818
-
RLAS-BIABC: A Reinforcement Learning-Based Answer Selection Using the BERT Model Boosted by an Improved ABC Algorithm 7 Jan 2023 · 0 repositories · arXiv:2301.02807
-
Critical Perspectives: A Benchmark Revealing Pitfalls in PerspectiveAPI 5 Jan 2023 · 1 repository · arXiv:2301.01874
-
Learning Trajectory-Word Alignments for Video-Language Tasks 5 Jan 2023 · 0 repositories · arXiv:2301.01953
-
Sequentially Controlled Text Generation 5 Jan 2023 · 0 repositories · arXiv:2301.02299
-
InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval 4 Jan 2023 · 1 repository · arXiv:2301.01820
-
UniHD at TSAR-2022 Shared Task: Is Compute All We Need for Lexical Simplification? 4 Jan 2023 · 1 repository · arXiv:2301.01764
-
Large Language Models as Corporate Lobbyists 3 Jan 2023 · 1 repository · arXiv:2301.01181
-
PIE-QG: Paraphrased Information Extraction for Unsupervised Question Generation from Small Corpora 3 Jan 2023 · 0 repositories · arXiv:2301.01064
-
Understanding Political Polarisation using Language Models: A dataset and method 2 Jan 2023 · 0 repositories · arXiv:2301.00891
-
Floods Relevancy and Identification of Location from Twitter Posts using NLP Techniques 1 Jan 2023 · 0 repositories · arXiv:2301.00321
-
Fusing Pre-Trained Language Models With Multimodal Prompts Through Reinforcement Learning 1 Jan 2023 · 1 repository
-
Generating Human Motion From Textual Descriptions With Discrete Representations 1 Jan 2023 · 0 repositories
-
Leveraging Semantic Representations Combined with Contextual Word Representations for Recognizing Textual Entailment in Vietnamese 1 Jan 2023 · 0 repositories · arXiv:2301.00422