Methods › General › Regularization › Weight Decay › Papers, page 71
Weight Decay
Papers archive 2025-07-28
archive papers tagged: 10,713 · with a code link: 4,533 · where Syntology ran a sample: 1,291 (1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,291 of 10,713 tagged: 1,064 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 71 of 108: papers 7,001 to 7,100 of 10,713, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Given Users Recommendations Based on Reviews on Yelp 3 Dec 2021 · 1 repository · arXiv:2112.01762
-
NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference 3 Dec 2021 · 0 repositories · arXiv:2112.02191
-
Siamese BERT-based Model for Web Search Relevance Ranking Evaluated on a New Czech Dataset 3 Dec 2021 · 1 repository · arXiv:2112.01810
-
BEVT: BERT Pretraining of Video Transformers 2 Dec 2021 · 1 repository · arXiv:2112.01529
-
PLSUM: Generating PT-BR Wikipedia by Summarizing Multiple Websites 2 Dec 2021 · 1 repository · arXiv:2112.01591
-
Unsupervised Law Article Mining based on Deep Pre-Trained Language Representation Models with Application to the Italian Civil Code 2 Dec 2021 · 0 repositories · arXiv:2112.03033
-
Combining Global and Local Attention with Positional Encoding for Video Summarization 1 Dec 2021 · 1 repository
-
Domain-oriented Language Pre-training with Adaptive Hybrid Masking and Optimal Transport Alignment 1 Dec 2021 · 0 repositories · arXiv:2112.03024
-
DRONE: Data-aware Low-rank Compression for Large NLP Models 1 Dec 2021 · 0 repositories
-
NER-BERT: A Pre-trained Model for Low-Resource Entity Tagging 1 Dec 2021 · 0 repositories · arXiv:2112.00405
-
Searching for Efficient Transformers for Language Modeling 1 Dec 2021 · 0 repositories
-
Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight Decay 1 Dec 2021 · 0 repositories
-
Think Big, Teach Small: Do Language Models Distil Occam’s Razor? 1 Dec 2021 · 1 repository
-
TriBERT: Human-centric Audio-visual Representation Learning 1 Dec 2021 · 1 repository
-
Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer 1 Dec 2021 · 1 repository
-
Wiki to Automotive: Understanding the Distribution Shift and its impact on Named Entity Recognition 1 Dec 2021 · 0 repositories · arXiv:2112.00283
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
Chemical Identification and Indexing in PubMed Articles via BERT and Text-to-Text Approaches 30 Nov 2021 · 0 repositories · arXiv:2111.15622
-
Generating Rich Product Descriptions for Conversational E-commerce Systems 30 Nov 2021 · 0 repositories · arXiv:2111.15298
-
KARL-Trans-NER: Knowledge Aware Representation Learning for Named Entity Recognition using Transformers 30 Nov 2021 · 0 repositories · arXiv:2111.15436
-
NLP Techniques for Water Quality Analysis in Social Media Content 30 Nov 2021 · 0 repositories · arXiv:2112.11441
-
Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models 30 Nov 2021 · 1 repository · arXiv:2112.00029Syntology official: harvested, nothing ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 6 harvested samples) · 5 pointer-only (licence)
-
Sentiment Analysis and Effect of COVID-19 Pandemic using College SubReddit Data 30 Nov 2021 · 1 repository · arXiv:2112.04351
-
SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Editing 30 Nov 2021 · 0 repositories · arXiv:2112.00180
-
Text classification problems via BERT embedding method and graph convolutional neural network 30 Nov 2021 · 0 repositories · arXiv:2111.15379
-
Customer Sentiment Analysis using Weak Supervision for Customer-Agent Chat 29 Nov 2021 · 0 repositories · arXiv:2111.14282
-
Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling 29 Nov 2021 · 3 repositories · arXiv:2111.14819Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Speech Tasks Relevant to Sleepiness Determined with Deep Transfer Learning 29 Nov 2021 · 0 repositories · arXiv:2111.14684
-
Context Matters in Semantically Controlled Language Generation for Task-oriented Dialogue Systems 28 Nov 2021 · 0 repositories · arXiv:2111.14119
-
Tapping BERT for Preposition Sense Disambiguation 27 Nov 2021 · 0 repositories · arXiv:2111.13972
-
Predicting Document Coverage for Relation Extraction 26 Nov 2021 · 0 repositories · arXiv:2111.13611
-
Domain Prompt Learning for Efficiently Adapting CLIP to Unseen Domains 25 Nov 2021 · 1 repository · arXiv:2111.12853Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Does constituency analysis enhance domain-specific pre-trained BERT models for relation extraction? 25 Nov 2021 · 0 repositories · arXiv:2112.02955
-
Evaluating the Robustness of Retrieval Pipelines with Query Variation Generators 25 Nov 2021 · 1 repository · arXiv:2111.13057
-
New Approaches to Long Document Summarization: Fourier Transform Based Attention in a Transformer Model 25 Nov 2021 · 0 repositories · arXiv:2111.15473
-
Probabilistic Impact Score Generation using Ktrain-BERT to Identify Hate Words from Twitter Discussions 25 Nov 2021 · 0 repositories · arXiv:2111.12939
-
Recommending Multiple Positive Citations for Manuscript via Content-Dependent Modeling and Multi-Positive Triplet 25 Nov 2021 · 0 repositories · arXiv:2111.12899
-
Transformer-based Korean Pretrained Language Models: A Survey on Three Years of Progress 25 Nov 2021 · 0 repositories · arXiv:2112.03014
-
PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers 24 Nov 2021 · 1 repository · arXiv:2111.12710
-
DABS: A Domain-Agnostic Benchmark for Self-Supervised Learning 23 Nov 2021 · 1 repository · arXiv:2111.12062Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Variational Learning for Unsupervised Knowledge Grounded Dialogs 23 Nov 2021 · 1 repository · arXiv:2112.00653
-
Can depth-adaptive BERT perform better on binary classification tasks 22 Nov 2021 · 0 repositories · arXiv:2111.10951
-
Does BERT look at sentiment lexicon? 19 Nov 2021 · 0 repositories · arXiv:2111.10100
-
Lexicon-based Methods vs. BERT for Text Sentiment Analysis 19 Nov 2021 · 0 repositories · arXiv:2111.10097
-
ClipCap: CLIP Prefix for Image Captioning 18 Nov 2021 · 4 repositories · arXiv:2111.09734Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing 18 Nov 2021 · 3 repositories · arXiv:2111.09543Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Dynamic-TinyBERT: Boost TinyBERT's Inference Efficiency by Dynamic Sequence Length 18 Nov 2021 · 0 repositories · arXiv:2111.09645
-
LAnoBERT: System Log Anomaly Detection based on BERT Masked Language Model 18 Nov 2021 · 0 repositories · arXiv:2111.09564
-
RoBERTuito: a pre-trained language model for social media text in Spanish 18 Nov 2021 · 1 repository · arXiv:2111.09453Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval 18 Nov 2021 · 1 repository · arXiv:2111.09852
-
Guiding Generative Language Models for Data Augmentation in Few-Shot Text Classification 17 Nov 2021 · 0 repositories · arXiv:2111.09064
-
A Comparative Study on Transfer Learning and Distance Metrics in Semantic Clustering over the COVID-19 Tweets 16 Nov 2021 · 0 repositories · arXiv:2111.08658
-
A Flexible Multi-Task Model for BERT Serving 16 Nov 2021 · 0 repositories
-
A Graph Enhanced BERT Model for Event Prediction 16 Nov 2021 · 0 repositories
-
A Sentence is Worth 128 Pseudo Tokens: A Semantic-Aware Contrastive Learning Framework for Sentence Embeddings 16 Nov 2021 · 0 repositories
-
A Structured Semantic Reinforcement method for Task-Oriented Dialogue 16 Nov 2021 · 0 repositories
-
Active Dialogue Simulation in Conversational Systems 16 Nov 2021 · 0 repositories
-
AdapLeR: Speeding up Inference by Adaptive Length Reduction 16 Nov 2021 · 0 repositories
-
An Information Theoretic Measurement of Topical Relevance in Learner Essays 16 Nov 2021 · 0 repositories
-
An Isotropy Analysis in the Multilingual BERT Embedding Space 16 Nov 2021 · 0 repositories
-
ANNA: Enhanced Language Representation for Question Answering 16 Nov 2021 · 0 repositories
-
Assessing the Coherence Modeling Capabilities of Pretrained Transformer-based Language Models 16 Nov 2021 · 0 repositories
-
Attention-based Multi-hypothesis Fusion for Speech Summarization 16 Nov 2021 · 2 repositories · arXiv:2111.08201
-
BERT got a Date: Introducing Transformers to Temporal Tagging 16 Nov 2021 · 0 repositories
-
BERT is Robust! A Case Against Synonym-Based Adversarial Examples in Text Classification 16 Nov 2021 · 0 repositories
-
BigFive: A Dataset of Coarse- and Fine-Grained Personality Characteristics 16 Nov 2021 · 0 repositories
-
BORT: Back and Denoising Reconstruction for End-to-End Task-Oriented Dialog 16 Nov 2021 · 0 repositories
-
Building Chinese Biomedical Language Models via Multi-Level Text Discrimination 16 Nov 2021 · 1 repository
-
Chinese Word Attention based on Valid Division of Sentence 16 Nov 2021 · 0 repositories
-
Contextualized Sensorimotor Norms: multi-dimensional measures of sensorimotor strength for ambiguous English words, in context 16 Nov 2021 · 0 repositories
-
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation 16 Nov 2021 · 0 repositories
-
Cross-domain Named Entity Recognition via Graph Matching 16 Nov 2021 · 0 repositories
-
CVSS-BERT: Explainable Natural Language Processing to Determine the Severity of a Computer Security Vulnerability from its Description 16 Nov 2021 · 1 repository · arXiv:2111.08510
-
Data Augmentation for Intent Classification with Generic Large Language Models 16 Nov 2021 · 0 repositories
-
Data Contamination: From Memorization to Exploitation 16 Nov 2021 · 0 repositories
-
DAWSON: Data Augmentation using Weak Supervision On Natural Language 16 Nov 2021 · 0 repositories
-
Deep-to-bottom Weights Decay: A Systemic Knowledge Review Learning Technique for Transformer Layers in Knowledge Distillation 16 Nov 2021 · 0 repositories
-
Discontinuous Constituency and BERT: A Case Study of Dutch 16 Nov 2021 · 0 repositories
-
ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language Models 16 Nov 2021 · 0 repositories
-
ELLE: Efficient Lifelong Pre-training for Emerging Data 16 Nov 2021 · 0 repositories
-
End-to-end Task-oriented Dialog Policy Learning based on Pre-trained Language Model 16 Nov 2021 · 0 repositories
-
ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention 16 Nov 2021 · 0 repositories
-
Event Detection via Derangement Question Answering 16 Nov 2021 · 0 repositories
-
EventBERT 16 Nov 2021 · 0 repositories
-
Explicit Modeling the Context for Chinese NER 16 Nov 2021 · 0 repositories
-
Exploring and Adapting Chinese GPT to Pinyin Input Method 16 Nov 2021 · 0 repositories
-
Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences 16 Nov 2021 · 0 repositories
-
Feature-rich Open-vocabulary Interpretable Neural Representations for All of the World’s 7000 Languages 16 Nov 2021 · 0 repositories
-
Feature Structure Distillation for BERT Transferring 16 Nov 2021 · 0 repositories
-
gaBERT — an Irish Language Model 16 Nov 2021 · 0 repositories
-
Generative Pre-Trained Transformer for Design Concept Generation: An Exploration 16 Nov 2021 · 0 repositories · arXiv:2111.08489
-
Get the Point! Graph Enhanced Candidate Retrieval for Zero-shot Entity Linking 16 Nov 2021 · 0 repositories
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 16 Nov 2021 · 0 repositories
-
Graph-based Fine-grained Multimodal Attention Mechanism for Sentiment Analysis 16 Nov 2021 · 0 repositories
-
How does the pre-training objective affect what large language models learn about linguistic properties? 16 Nov 2021 · 0 repositories
-
Impact of Tokenization on Language Models: An Analysis for Turkish 16 Nov 2021 · 0 repositories
-
Improving GPT-3 after deployment with a dynamic memory of feedback 16 Nov 2021 · 0 repositories
-
Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation 16 Nov 2021 · 0 repositories
-
Improving Unsupervised Sentence Simplification Using Fine-Tuned Masked Language Models 16 Nov 2021 · 0 repositories
-
Input-specific Attention Subnetworks for Adversarial Detection 16 Nov 2021 · 0 repositories