Methods › General › Regularization › Weight Decay › Papers, page 63
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 63 of 108: papers 6,201 to 6,300 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond 30 Jul 2022 · 0 repositories · arXiv:2208.00173
-
Code Comment Inconsistency Detection with BERT and Longformer 29 Jul 2022 · 1 repository · arXiv:2207.14444
-
Curriculum Learning for Data-Efficient Vision-Language Alignment 29 Jul 2022 · 0 repositories · arXiv:2207.14525
-
SERCNN: Stacked Embedding Recurrent Convolutional Neural Network in Detecting Depression on Twitter 29 Jul 2022 · 0 repositories · arXiv:2207.14535
-
CrAM: A Compression-Aware Minimizer 28 Jul 2022 · 1 repository · arXiv:2207.14200Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
LAD: Language Models as Data for Zero-Shot Dialog 28 Jul 2022 · 0 repositories · arXiv:2207.14393
-
Large Language Models and the Reverse Turing Test 28 Jul 2022 · 0 repositories · arXiv:2207.14382
-
SDBERT: SparseDistilBERT, a faster and smaller BERT model 28 Jul 2022 · 0 repositories · arXiv:2208.10246
-
Sequence to sequence pretraining for a less-resourced Slovenian language 28 Jul 2022 · 1 repository · arXiv:2207.13988
-
Distributional Actor-Critic Ensemble for Uncertainty-Aware Continuous Control 27 Jul 2022 · 0 repositories · arXiv:2207.13730
-
SoundChoice: Grapheme-to-Phoneme Models with Semantic Disambiguation 27 Jul 2022 · 1 repository · arXiv:2207.13703
-
Bundle MCR: Towards Conversational Bundle Recommendation 26 Jul 2022 · 1 repository · arXiv:2207.12628
-
Fine-Tuning BERT for Automatic ADME Semantic Labeling in FDA Drug Labeling to Enhance Product-Specific Guidance Assessment 25 Jul 2022 · 0 repositories · arXiv:2207.12376
-
Is GPT-3 all you need for Visual Question Answering in Cultural Heritage? 25 Jul 2022 · 0 repositories · arXiv:2207.12101
-
A Cognitive Study on Semantic Similarity Analysis of Large Corpora: A Transformer-based Approach 24 Jul 2022 · 0 repositories · arXiv:2207.11716
-
Better Reasoning Behind Classification Predictions with BERT for Fake News Detection 23 Jul 2022 · 0 repositories · arXiv:2207.11562
-
Zero-Shot Video Captioning with Evolving Pseudo-Tokens 22 Jul 2022 · 1 repository · arXiv:2207.11100
-
BigIssue: A Realistic Bug Localization Benchmark 21 Jul 2022 · 0 repositories · arXiv:2207.10739
-
Efficient model compression with Random Operation Access Specific Tile (ROAST) hashing 21 Jul 2022 · 1 repository · arXiv:2207.10702
-
Abstract Demonstrations and Adaptive Exploration for Efficient and Stable Multi-step Sparse Reward Reinforcement Learning 19 Jul 2022 · 1 repository · arXiv:2207.09243
-
Enhancing Collaborative Filtering Recommender with Prompt-Based Sentiment Analysis 19 Jul 2022 · 1 repository · arXiv:2207.12883
-
PiC: A Phrase-in-Context Dataset for Phrase Understanding and Semantic Search 19 Jul 2022 · 1 repository · arXiv:2207.09068
-
Pre-trained language models with domain knowledge for biomedical extractive summarization 19 Jul 2022 · 1 repository
-
Revealing Secrets From Pre-trained Models 19 Jul 2022 · 0 repositories · arXiv:2207.09539
-
Selection Bias Induced Spurious Correlations in Large Language Models 18 Jul 2022 · 1 repository · arXiv:2207.08982
-
Word Play for Playing Othello (Reverses) 18 Jul 2022 · 0 repositories · arXiv:2207.08766
-
Aspect-specific Context Modeling for Aspect-based Sentiment Analysis 17 Jul 2022 · 1 repository · arXiv:2207.08099
-
Can large language models reason about medical questions? 17 Jul 2022 · 1 repository · arXiv:2207.08143Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
ELECTRA is a Zero-Shot Learner, Too 17 Jul 2022 · 1 repository · arXiv:2207.08141Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Representation Learning of Image Schema 17 Jul 2022 · 0 repositories · arXiv:2207.08256
-
Robust Action Governor for Uncertain Piecewise Affine Systems with Non-convex Constraints and Safe Reinforcement Learning 17 Jul 2022 · 0 repositories · arXiv:2207.08240
-
A Context-Sensitive Word Embedding Approach for The Detection of Troll Tweets 17 Jul 2022 · 0 repositories · arXiv:2207.08230
-
POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging 15 Jul 2022 · 1 repository · arXiv:2207.07697
-
Position Prediction as an Effective Pretraining Strategy 15 Jul 2022 · 1 repository · arXiv:2207.07611
-
Z-Index at CheckThat! Lab 2022: Check-Worthiness Identification on Tweet Text 15 Jul 2022 · 0 repositories · arXiv:2207.07308
-
Combing for Credentials: Active Pattern Extraction from Smart Reply 14 Jul 2022 · 0 repositories · arXiv:2207.10802
-
Bootstrapped Masked Autoencoders for Vision BERT Pretraining 14 Jul 2022 · 1 repository · arXiv:2207.07116Syntology official (archive's flag): 6 ran · 6 ran (of which 6 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 6 samples that ran constructed an object rather than computing a result (of 7 harvested samples) · 7 pointer-only (licence)
-
Language Modelling with Pixels 14 Jul 2022 · 1 repository · arXiv:2207.06991Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Multilinguals at SemEval-2022 Task 11: Complex NER in Semantically Ambiguous Settings for Low Resource Languages 14 Jul 2022 · 1 repository · arXiv:2207.06882
-
PIAT: Physics Informed Adversarial Training for Solving Partial Differential Equations 14 Jul 2022 · 1 repository · arXiv:2207.06647
-
A Transfer Learning Based Model for Text Readability Assessment in German 13 Jul 2022 · 0 repositories · arXiv:2207.06265
-
DynaST: Dynamic Sparse Transformer for Exemplar-Guided Image Generation 13 Jul 2022 · 1 repository · arXiv:2207.06124Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Exploiting Word Semantics to Enrich Character Representations of Chinese Pre-trained Models 13 Jul 2022 · 1 repository · arXiv:2207.05928
-
Re2G: Retrieve, Rerank, Generate 13 Jul 2022 · 1 repository · arXiv:2207.06300Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
How Do Multilingual Encoders Learn Cross-lingual Representation? 12 Jul 2022 · 0 repositories · arXiv:2207.05737
-
Using Paraphrases to Study Properties of Contextual Embeddings 12 Jul 2022 · 0 repositories · arXiv:2207.05553
-
Learning Large-scale Universal User Representation with Sparse Mixture of Experts 11 Jul 2022 · 0 repositories · arXiv:2207.04648
-
Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition 11 Jul 2022 · 1 repository · arXiv:2207.04697
-
Overview of the Shared Task on Fake News Detection in Urdu at FIRE 2021 11 Jul 2022 · 0 repositories · arXiv:2207.05133
-
SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning 11 Jul 2022 · 2 repositories · arXiv:2207.04606Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 2 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Multilingual Persuasion Detection: Video Games as an Invaluable Data Source for NLP 10 Jul 2022 · 1 repository · arXiv:2207.04453
-
Few-shot training LLMs for project-specific code-summarization 9 Jul 2022 · 0 repositories · arXiv:2207.04237
-
Computationally Identifying Funneling and Focusing Questions in Classroom Discourse 8 Jul 2022 · 1 repository · arXiv:2208.04715
-
Deep Visual-Linguistic Fusion Network Considering Cross-Modal Inconsistency for Rumor Detection 8 Jul 2022 · 3 repositories
-
Hidden Schema Networks 8 Jul 2022 · 0 repositories · arXiv:2207.03777
-
A Large Scale Search Dataset for Unbiased Learning to Rank 7 Jul 2022 · 1 repository · arXiv:2207.03051
-
Active Learning and Multi-label Classification for Ellipsis and Coreference Detection in Conversational Question-Answering 7 Jul 2022 · 0 repositories · arXiv:2207.03145
-
AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition 7 Jul 2022 · 0 repositories · arXiv:2207.03422
-
Neural Language Models are not Born Equal to Fit Brain Data, but Training Helps 7 Jul 2022 · 0 repositories · arXiv:2207.03380
-
Sensitivity Analysis on Transferred Neural Architectures of BERT and GPT-2 for Financial Sentiment Analysis 7 Jul 2022 · 0 repositories · arXiv:2207.03037
-
Ask Me What You Need: Product Retrieval using Knowledge from GPT-3 6 Jul 2022 · 0 repositories · arXiv:2207.02516
-
Aspect-Based Sentiment Analysis using Local Context Focus Mechanism with DeBERTa 6 Jul 2022 · 0 repositories · arXiv:2207.02424
-
SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval 6 Jul 2022 · 1 repository · arXiv:2207.02578
-
The Role of Complex NLP in Transformers for Text Ranking? 6 Jul 2022 · 0 repositories · arXiv:2207.02522
-
Betti numbers of attention graphs is all you really need 5 Jul 2022 · 1 repository · arXiv:2207.01903
-
Machine Learning Model Sizes and the Parameter Gap 5 Jul 2022 · 0 repositories · arXiv:2207.02852
-
BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model 4 Jul 2022 · 0 repositories · arXiv:2207.01718
-
Using contextual sentence analysis models to recognize ESG concepts 4 Jul 2022 · 0 repositories · arXiv:2207.01402
-
GUIM -- General User and Item Embedding with Mixture of Representation in E-commerce 2 Jul 2022 · 0 repositories · arXiv:2207.00750
-
A Polyphone BERT for Polyphone Disambiguation in Mandarin Chinese 1 Jul 2022 · 0 repositories · arXiv:2207.12089
-
AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection 30 Jun 2022 · 1 repository · arXiv:2206.15476Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Compressing Pre-trained Transformers via Low-Bit NxM Sparsity for Natural Language Understanding 30 Jun 2022 · 0 repositories · arXiv:2206.15014
-
DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale 30 Jun 2022 · 2 repositories · arXiv:2207.00032
-
GaitForeMer: Self-Supervised Pre-Training of Transformers via Human Motion Forecasting for Few-Shot Gait Impairment Severity Estimation 30 Jun 2022 · 1 repository · arXiv:2207.00106
-
ListBERT: Learning to Rank E-commerce products with Listwise BERT 30 Jun 2022 · 0 repositories · arXiv:2206.15198
-
The Topological BERT: Transforming Attention into Topology for Natural Language Processing 30 Jun 2022 · 0 repositories · arXiv:2206.15195
-
Two-Stage Classifier for COVID-19 Misinformation Detection Using BERT: a Study on Indonesian Tweets 30 Jun 2022 · 2 repositories · arXiv:2206.15359
-
Chinese Word Sense Embedding with SememeWSD and Synonym Set 29 Jun 2022 · 2 repositories · arXiv:2206.14388
-
SALO: An Efficient Spatial Accelerator Enabling Hybrid Sparse Attention Mechanisms for Long Sequences 29 Jun 2022 · 0 repositories · arXiv:2206.14550
-
Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody 29 Jun 2022 · 0 repositories · arXiv:2206.14643
-
Summarizing Videos using Concentrated Attention and Considering the Uniqueness and Diversity of the Video Frames 29 Jun 2022 · 1 repository
-
Two-Stage COVID19 Classification Using BERT Features 29 Jun 2022 · 0 repositories · arXiv:2206.14861
-
Exploring linguistic feature and model combination for speech recognition based automatic AD detection 28 Jun 2022 · 0 repositories · arXiv:2206.13758
-
Materials Transformers Language Models for Generative Materials Design: a benchmark study 27 Jun 2022 · 1 repository · arXiv:2206.13578
-
A multi-model-based deep learning framework for short text multiclass classification with the imbalanced and extremely small data set 24 Jun 2022 · 0 repositories · arXiv:2206.12027
-
A Test for Evaluating Performance in Human-Computer Systems 24 Jun 2022 · 0 repositories · arXiv:2206.12390
-
Text and author-level political inference using heterogeneous knowledge representations 24 Jun 2022 · 0 repositories · arXiv:2206.12293
-
Unified BERT for Few-shot Natural Language Understanding 24 Jun 2022 · 0 repositories · arXiv:2206.12094
-
Using BERT Embeddings to Model Word Importance in Conversational Transcripts for Deaf and Hard of Hearing Users 24 Jun 2022 · 0 repositories · arXiv:2206.12368
-
A Disability Lens towards Biases in GPT-3 Generated Open-Ended Languages 23 Jun 2022 · 0 repositories · arXiv:2206.11993
-
BERT Rankers are Brittle: a Study using Adversarial Document Perturbations 23 Jun 2022 · 1 repository · arXiv:2206.11724
-
Revisiting Orthogonality Regularization: A Study for Convolutional Neural Networks in Image Classification 23 Jun 2022 · 1 repository
-
Towards WinoQueer: Developing a Benchmark for Anti-Queer Bias in Large Language Models 23 Jun 2022 · 0 repositories · arXiv:2206.11484
-
Answer Fast: Accelerating BERT on the Tensor Streaming Processor 22 Jun 2022 · 0 repositories · arXiv:2206.11062
-
Efficient and effective training of language and graph neural network models 22 Jun 2022 · 0 repositories · arXiv:2206.10781
-
An Automatic and Efficient BERT Pruning for Edge AI Systems 21 Jun 2022 · 0 repositories · arXiv:2206.10461
-
CoCoPIE XGen: A Full-Stack AI-Oriented Optimizing Framework 21 Jun 2022 · 0 repositories · arXiv:2206.10620
-
Knowledge Graph Fusion for Language Model Fine-tuning 21 Jun 2022 · 0 repositories · arXiv:2206.14574
-
TAPHSIR: Towards AnaPHoric Ambiguity Detection and ReSolution In Requirements 21 Jun 2022 · 0 repositories · arXiv:2206.10227
-
Using cognitive psychology to understand GPT-3 21 Jun 2022 · 0 repositories · arXiv:2206.14576