Methods › General › Regularization › Label Smoothing › Papers, page 118
Label Smoothing
Papers archive 2025-07-28
archive papers tagged: 14,327 · with a code link: 6,651 · where Syntology ran a sample: 2,259 (1,920 with a run with no instrument failure, 339 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,259 of 14,327 tagged: 1,920 with a run with no instrument failure, 339 where every run was a failure of Syntology's instrument)
Page 118 of 144: papers 11,701 to 11,800 of 14,327, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Guiding Transformers to Process in Steps 29 Sep 2021 · 0 repositories
-
HFSP: A Hardware-friendly Soft Pruning Framework for Vision Transformers 29 Sep 2021 · 0 repositories
-
Hierarchical Character Tagger for Short Text Spelling Error Correction 29 Sep 2021 · 0 repositories · arXiv:2109.14259
-
HoloFormer: Deep Compression of Pre-Trained Transforms via Unified Optimization of N:M Sparsity and Integer Quantization 29 Sep 2021 · 0 repositories
-
Isotropic Contextual Representations through Variational Regularization 29 Sep 2021 · 0 repositories
-
Learning to Schedule Learning rate with Graph Neural Networks 29 Sep 2021 · 0 repositories
-
LMSA: Low-relation Mutil-head Self-Attention Mechanism in Visual Transformer 29 Sep 2021 · 0 repositories
-
MaiT: integrating spatial locality into image transformers with attention masks 29 Sep 2021 · 1 repository
-
MLP-based architecture with variable length input for automatic speech recognition 29 Sep 2021 · 0 repositories
-
Non-Autoregressive Models are Better Multilingual Translators 29 Sep 2021 · 0 repositories
-
Offline Pre-trained Multi-Agent Decision Transformer 29 Sep 2021 · 0 repositories
-
Pretraining for Language Conditioned Imitation with Transformers 29 Sep 2021 · 0 repositories
-
Privacy-preserving Task-Agnostic Vision Transformer for Image Processing 29 Sep 2021 · 1 repository
-
Pseudo Knowledge Distillation: Towards Learning Optimal Instance-specific Label Smoothing Regularization 29 Sep 2021 · 0 repositories
-
Scale Efficiently: Insights from Pretraining and Finetuning Transformers 29 Sep 2021 · 0 repositories
-
Scaling the Depth of Vision Transformers via the Fourier Domain Analysis 29 Sep 2021 · 0 repositories
-
Semi-supervised Offline Reinforcement Learning with Pre-trained Decision Transformers 29 Sep 2021 · 0 repositories
-
SiT: Simulation Transformer for Particle-based Physics Simulation 29 Sep 2021 · 0 repositories
-
Spanning Tree-based Graph Generation for Molecules 29 Sep 2021 · 0 repositories
-
Sparse Attention with Learning to Hash 29 Sep 2021 · 0 repositories
-
Specialized Transformers: Faster, Smaller and more Accurate NLP Models 29 Sep 2021 · 0 repositories
-
Subdimensional Expansion Using Attention-Based Learning For Multi-Agent Path Finding 29 Sep 2021 · 1 repository · arXiv:2109.14695
-
Temporal Action Localization with Global Segmentation Mask Transformers 29 Sep 2021 · 0 repositories
-
To Smooth or not to Smooth? On Compatibility between Label Smoothing and Knowledge Distillation 29 Sep 2021 · 0 repositories
-
Topic Aware Neural Language Model: Domain Adaptation of Unconditional Text Generation Models 29 Sep 2021 · 0 repositories
-
TransTCN: An Attention-based TCN Framework for Sequential Modeling 29 Sep 2021 · 0 repositories
-
Tuformer: Data-Driven Design of Expressive Transformer by Tucker Tensor Representation 29 Sep 2021 · 0 repositories
-
UFO-ViT: High Performance Linear Vision Transformer without Softmax 29 Sep 2021 · 1 repository · arXiv:2109.14382
-
Understanding Generalized Label Smoothing when Learning with Noisy Labels 29 Sep 2021 · 1 repository
-
Understanding the robustness-accuracy tradeoff by rethinking robust fairness 29 Sep 2021 · 0 repositories
-
Understanding the Role of Self Attention for Efficient Speech Recognition 29 Sep 2021 · 0 repositories
-
Video Forgery Detection Using Multiple Cues on Fusion of EfficientNet and Swin Transformer 29 Sep 2021 · 0 repositories
-
VUT: Versatile UI Transformer for Multimodal Multi-Task User Interface Modeling 29 Sep 2021 · 0 repositories
-
Fine-tuning Vision Transformers for the Prediction of State Variables in Ising Models 28 Sep 2021 · 0 repositories · arXiv:2109.13925
-
Nana-HDR: A Non-attentive Non-autoregressive Hybrid Model for TTS 28 Sep 2021 · 0 repositories · arXiv:2109.13673
-
Single-dataset Experts for Multi-dataset Question Answering 28 Sep 2021 · 1 repository · arXiv:2109.13880Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates 27 Sep 2021 · 1 repository · arXiv:2109.12804
-
Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information 27 Sep 2021 · 1 repository · arXiv:2109.13073
-
Improving Uncertainty of Deep Learning-based Object Classification on Radar Spectra using Label Smoothing 27 Sep 2021 · 0 repositories · arXiv:2109.12851
-
Integrated Training for Sequence-to-Sequence Models Using Non-Autoregressive Transformer 27 Sep 2021 · 0 repositories · arXiv:2109.12950
-
Sparse Spatial Transformers for Few-Shot Learning 27 Sep 2021 · 1 repository · arXiv:2109.12932Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Multi-Transformer: A New Neural Network-Based Architecture for Forecasting S&P Volatility 26 Sep 2021 · 1 repository · arXiv:2109.12621
-
Vision Transformer Hashing for Image Retrieval 26 Sep 2021 · 1 repository · arXiv:2109.12564
-
ViT Cane: Visual Assistant for the Visually Impaired 26 Sep 2021 · 0 repositories · arXiv:2109.13857
-
A real-time and high-precision method for small traffic-signs recognition 25 Sep 2021 · 1 repository
-
TEMGNet: Deep Transformer-based Decoding of Upperlimb sEMG for Hand Gestures Recognition 25 Sep 2021 · 0 repositories · arXiv:2109.12379
-
DACT-BERT: Differentiable Adaptive Computation Time for an Efficient BERT Inference 24 Sep 2021 · 0 repositories · arXiv:2109.11745
-
Identification of Enzymatic Active Sites with Unsupervised Language Modeling 24 Sep 2021 · 0 repositories
-
Localizing Infinity-shaped fishes: Sketch-guided object localization in the wild 24 Sep 2021 · 0 repositories · arXiv:2109.11874
-
Long-Range Transformers for Dynamic Spatiotemporal Forecasting 24 Sep 2021 · 2 repositories · arXiv:2109.12218Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Training dataset generation for bridge game registration 24 Sep 2021 · 0 repositories · arXiv:2109.11861
-
Transformers Generalize Linearly 24 Sep 2021 · 1 repository · arXiv:2109.12036
-
Dependency Structure for News Document Summarization 23 Sep 2021 · 0 repositories · arXiv:2109.11199
-
OH-Former: Omni-Relational High-Order Transformer for Person Re-Identification 23 Sep 2021 · 0 repositories · arXiv:2109.11159
-
The Volctrans GLAT System: Non-autoregressive Translation Meets WMT21 23 Sep 2021 · 0 repositories · arXiv:2109.11247
-
Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models 22 Sep 2021 · 1 repository · arXiv:2109.11058
-
Hierarchical Multimodal Transformer to Summarize Videos 22 Sep 2021 · 0 repositories · arXiv:2109.10559
-
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation 22 Sep 2021 · 1 repository · arXiv:2109.10504
-
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers 22 Sep 2021 · 3 repositories · arXiv:2109.10686Syntology community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 4 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
T6D-Direct: Transformers for Multi-Object 6D Pose Direct Regression 22 Sep 2021 · 0 repositories · arXiv:2109.10948
-
The NiuTrans Machine Translation Systems for WMT21 22 Sep 2021 · 0 repositories · arXiv:2109.10485
-
DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers 21 Sep 2021 · 1 repository · arXiv:2109.10060
-
LOTR: Face Landmark Localization Using Localization Transformer 21 Sep 2021 · 0 repositories · arXiv:2109.10057
-
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models 21 Sep 2021 · 8 repositories · arXiv:2109.10282Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Dyadformer: A Multi-modal Transformer for Long-Range Modeling of Dyadic Interactions 20 Sep 2021 · 0 repositories · arXiv:2109.09487
-
Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends 20 Sep 2021 · 1 repository · arXiv:2109.09824Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Do Long-Range Language Models Actually Use Long-Range Context? 19 Sep 2021 · 0 repositories · arXiv:2109.09115
-
The Seismo-Performer: A Novel Machine Learning Approach for General and Efficient Seismic Phase Recognition from Local Earthquakes in Real Time 19 Sep 2021 · 1 repository
-
UNetFormer: A UNet-like Transformer for Efficient Semantic Segmentation of Remote Sensing Urban Scene Imagery 18 Sep 2021 · 1 repository · arXiv:2109.08937Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
SDTP: Semantic-aware Decoupled Transformer Pyramid for Dense Image Prediction 18 Sep 2021 · 0 repositories · arXiv:2109.08963
-
Towards High-Quality Temporal Action Detection with Sparse Proposals 18 Sep 2021 · 1 repository · arXiv:2109.08847
-
Continuous Streaming Multi-Talker ASR with Dual-path Transducers 17 Sep 2021 · 0 repositories · arXiv:2109.08555
-
Digging Errors in NMT: Evaluating and Understanding Model Errors from Hypothesis Distribution 17 Sep 2021 · 0 repositories
-
Expression Snippet Transformer for Robust Video-based Facial Expression Recognition 17 Sep 2021 · 0 repositories · arXiv:2109.08409
-
Focus on the Target's Vocabulary: Masked Label Smoothing for Machine Translation 17 Sep 2021 · 0 repositories
-
From Known to Unknown: Knowledge-guided Transformer for Time-Series Sales Forecasting in Alibaba 17 Sep 2021 · 0 repositories · arXiv:2109.08381
-
Learning Low-frequency Patterns with A Pre-trained Document-Grounded Conversation Model 17 Sep 2021 · 0 repositories
-
Primer: Searching for Efficient Transformers for Language Modeling 17 Sep 2021 · 4 repositories · arXiv:2109.08668Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling? An Extensive Empirical Study on Language Tasks 17 Sep 2021 · 0 repositories
-
The JHU-Microsoft Submission for WMT21 Quality Estimation Shared Task 17 Sep 2021 · 0 repositories · arXiv:2109.08724
-
An End-to-End Transformer Model for 3D Object Detection 16 Sep 2021 · 1 repository · arXiv:2109.08141Syntology 6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Fast-Slow Transformer for Visually Grounding Speech 16 Sep 2021 · 1 repository · arXiv:2109.08186
-
Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning 16 Sep 2021 · 1 repository · arXiv:2109.07799
-
MeLT: Message-Level Transformer with Masked Document Representations as Pre-Training for Stance Detection 16 Sep 2021 · 1 repository · arXiv:2109.08113
-
Scaling Laws for Neural Machine Translation 16 Sep 2021 · 0 repositories · arXiv:2109.07740
-
Sparse Factorization of Large Square Matrices 16 Sep 2021 · 1 repository · arXiv:2109.08184
-
TANet: A new Paradigm for Global Face Super-resolution via Transformer-CNN Aggregation Network 16 Sep 2021 · 0 repositories · arXiv:2109.08174
-
The NiuTrans System for the WMT21 Efficiency Task 16 Sep 2021 · 1 repository · arXiv:2109.08003
-
The NiuTrans System for WNGT 2020 Efficiency Task 16 Sep 2021 · 2 repositories · arXiv:2109.08008Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Utterance-level neural confidence measure for end-to-end children speech recognition 16 Sep 2021 · 0 repositories · arXiv:2109.07750
-
Anchor DETR: Query Design for Transformer-Based Object Detection 15 Sep 2021 · 2 repositories · arXiv:2109.07107
-
Complementary Feature Enhanced Network with Vision Transformer for Image Dehazing 15 Sep 2021 · 1 repository · arXiv:2109.07100
-
Incorporating Residual and Normalization Layers into Analysis of Masked Language Models 15 Sep 2021 · 2 repositories · arXiv:2109.07152Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MISSFormer: An Effective Medical Image Segmentation Transformer 15 Sep 2021 · 1 repository · arXiv:2109.07162
-
PnP-DETR: Towards Efficient Visual Analysis with Transformers 15 Sep 2021 · 1 repository · arXiv:2109.07036Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Pose Transformers (POTR): Human Motion Prediction with Non-Autoregressive Transformers 15 Sep 2021 · 1 repository · arXiv:2109.07531
-
RetroPrime: A Diverse, plausible and Transformer-based method for Single-Step retrosynthesis predictions 15 Sep 2021 · 1 repository
-
Sequence Length is a Domain: Length-based Overfitting in Transformer Models 15 Sep 2021 · 1 repository · arXiv:2109.07276
-
SupCL-Seq: Supervised Contrastive Learning for Downstream Optimized Sequence Representations 15 Sep 2021 · 1 repository · arXiv:2109.07424
-
Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU 15 Sep 2021 · 1 repository · arXiv:2109.07364