Methods › General › Regularization › Label Smoothing › Papers, page 137
Label Smoothing
Papers archive 2025-07-28
archive papers tagged: 14,327 · with a code link: 6,651 · where Syntology ran a sample: 2,259 (1,920 with a run with no instrument failure, 339 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,259 of 14,327 tagged: 1,920 with a run with no instrument failure, 339 where every run was a failure of Syntology's instrument)
Page 137 of 144: papers 13,601 to 13,700 of 14,327, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Insertion-Deletion Transformer 15 Jan 2020 · 0 repositories · arXiv:2001.05540
-
Non-Autoregressive Machine Translation with Disentangled Context Transformer 15 Jan 2020 · 1 repository · arXiv:2001.05136
-
Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture 15 Jan 2020 · 0 repositories · arXiv:2001.08290
-
Auto Completion of User Interface Layout Design Using Transformer-Based Tree Decoders 14 Jan 2020 · 0 repositories · arXiv:2001.05308
-
The problems with using STNs to align CNN feature maps 14 Jan 2020 · 0 repositories · arXiv:2001.05858
-
Reformer: The Efficient Transformer 13 Jan 2020 · 10 repositories · arXiv:2001.04451Syntology community repositories only · 6 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Attribute-guided Feature Learning Network for Vehicle Re-identification 12 Jan 2020 · 0 repositories · arXiv:2001.03872
-
Urdu-English Machine Transliteration using Neural Networks 12 Jan 2020 · 0 repositories · arXiv:2001.05296
-
Spatial-Temporal Transformer Networks for Traffic Flow Forecasting 9 Jan 2020 · 1 repository · arXiv:2001.02908
-
Streaming automatic speech recognition with the transformer model 8 Jan 2020 · 0 repositories · arXiv:2001.02674
-
RECAST: Interactive Auditing of Automatic Toxicity Detection Models 7 Jan 2020 · 0 repositories · arXiv:2001.01819
-
Regularization via Structural Label Smoothing 7 Jan 2020 · 0 repositories · arXiv:2001.01900
-
FDFtNet: Facing Off Fake Images using Fake Detection Fine-tuning Network 5 Jan 2020 · 2 repositories · arXiv:2001.01265Syntology official (archive's flag): 4 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples)
-
Learning Accurate Integer Transformer Machine-Translation Models 3 Jan 2020 · 0 repositories · arXiv:2001.00926
-
Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation 3 Jan 2020 · 1 repository · arXiv:2001.00891
-
Representing Unordered Data Using Complex-Weighted Multiset Automata 2 Jan 2020 · 0 repositories · arXiv:2001.00610
-
Attacking Lifelong Learning Models with Gradient Reversion 1 Jan 2020 · 0 repositories
-
BERT-AL: BERT for Arbitrarily Long Document Understanding 1 Jan 2020 · 0 repositories
-
NEURAL EXECUTION ENGINES 1 Jan 2020 · 0 repositories
-
Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring 1 Jan 2020 · 2 repositories
-
ZeroQ: A Novel Zero Shot Quantization Framework 1 Jan 2020 · 3 repositories · arXiv:2001.00281Syntology community repositories only · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 12 unverified (of 19 harvested samples) · 2 pointer-only (licence)
-
Deep Attentive Ranking Networks for Learning to Order Sentences 31 Dec 2019 · 0 repositories · arXiv:2001.00056
-
EEG based Continuous Speech Recognition using Transformers 31 Dec 2019 · 0 repositories · arXiv:2001.00501
-
AraNet: A Deep Learning Toolkit for Arabic Social Media 30 Dec 2019 · 1 repository · arXiv:1912.13072
-
All-in-One Image-Grounded Conversational Agents 28 Dec 2019 · 0 repositories · arXiv:1912.12394
-
Encoding word order in complex embeddings 27 Dec 2019 · 1 repository · arXiv:1912.12333
-
Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention 27 Dec 2019 · 1 repository · arXiv:1912.11959
-
Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection 25 Dec 2019 · 2 repositories · arXiv:1912.11637Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 1 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Improving Abstractive Text Summarization with History Aggregation 24 Dec 2019 · 0 repositories · arXiv:1912.11046
-
Multi-Graph Transformer for Free-Hand Sketch Recognition 24 Dec 2019 · 1 repository · arXiv:1912.11258Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
end-to-end training of a large vocabulary end-to-end speech recognition system 22 Dec 2019 · 0 repositories · arXiv:1912.11040
-
Learning and Evaluating Contextual Embedding of Source Code 21 Dec 2019 · 2 repositories · arXiv:2001.00059
-
Are Transformers universal approximators of sequence-to-sequence functions? 20 Dec 2019 · 0 repositories · arXiv:1912.10077
-
Axial Attention in Multidimensional Transformers 20 Dec 2019 · 3 repositories · arXiv:1912.12180Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
ET-USB: Transformer-Based Sequential Behavior Modeling for Inbound Customer Service 20 Dec 2019 · 0 repositories · arXiv:1912.10852
-
Shareable Representations for Search Query Understanding 20 Dec 2019 · 0 repositories · arXiv:2001.04345
-
Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting 19 Dec 2019 · 36 repositories · arXiv:1912.09363Syntology 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
Meshed-Memory Transformer for Image Captioning 17 Dec 2019 · 2 repositories · arXiv:1912.08226Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
BERTQA -- Attention on Steroids 14 Dec 2019 · 0 repositories · arXiv:1912.10435
-
Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining 14 Dec 2019 · 2 repositories · arXiv:1912.06813
-
WaLDORf: Wasteless Language-model Distillation On Reading-comprehension 13 Dec 2019 · 0 repositories · arXiv:1912.06638
-
Linear Mode Connectivity and the Lottery Ticket Hypothesis 11 Dec 2019 · 2 repositories · arXiv:1912.05671
-
Encoding Musical Style with Transformer Autoencoders 10 Dec 2019 · 0 repositories · arXiv:1912.05537
-
Learning a Layout Transfer Network for Context Aware Object Detection 9 Dec 2019 · 0 repositories · arXiv:1912.03865
-
Transformer Based Reinforcement Learning For Games 9 Dec 2019 · 0 repositories · arXiv:1912.03918
-
Bidirectional Scene Text Recognition with a Single Decoder 8 Dec 2019 · 1 repository · arXiv:1912.03656
-
Personalized Patent Claim Generation and Measurement 7 Dec 2019 · 0 repositories · arXiv:1912.03502
-
Synchronous Transformers for End-to-End Speech Recognition 6 Dec 2019 · 0 repositories · arXiv:1912.02958
-
Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks 6 Dec 2019 · 0 repositories · arXiv:1912.03063
-
Scratch that! An Evolution-based Adversarial Attack against Neural Networks 5 Dec 2019 · 1 repository · arXiv:1912.02316
-
Self-Supervised Contextual Language Representation of Radiology Reports to Improve the Identification of Communication Urgency 5 Dec 2019 · 0 repositories · arXiv:1912.02703
-
AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue 4 Dec 2019 · 0 repositories · arXiv:1912.10160
-
TU Wien @ TREC Deep Learning '19 -- Simple Contextualization for Re-ranking 3 Dec 2019 · 1 repository · arXiv:1912.01385
-
Viewpoint-Aware Loss with Angular Regularization for Person Re-Identification 3 Dec 2019 · 1 repository · arXiv:1912.01300
-
Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events 2 Dec 2019 · 0 repositories · arXiv:1912.02615
-
BLiMP: The Benchmark of Linguistic Minimal Pairs for English 2 Dec 2019 · 4 repositories · arXiv:1912.00582Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Long Distance Relationships without Time Travel: Boosting the Performance of a Sparse Predictive Autoencoder in Sequence Modeling 2 Dec 2019 · 1 repository · arXiv:1912.01116
-
Multi-Scale Self-Attention for Text Classification 2 Dec 2019 · 0 repositories · arXiv:1912.00544
-
Neural Academic Paper Generation 2 Dec 2019 · 1 repository · arXiv:1912.01982
-
Solving Arithmetic Word Problems Automatically Using Transformer and Unambiguous Representations 2 Dec 2019 · 1 repository · arXiv:1912.00871
-
Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks 1 Dec 2019 · 0 repositories
-
Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition 28 Nov 2019 · 0 repositories · arXiv:1911.12487
-
DeFINE: DEep Factorized INput Token Embeddings for Neural Sequence Modeling 27 Nov 2019 · 1 repository · arXiv:1911.12385
-
Do Attention Heads in BERT Track Syntactic Dependencies? 27 Nov 2019 · 1 repository · arXiv:1911.12246
-
SimpleBooks: Long-term dependency book dataset with simplified English vocabulary for word-level language modeling 27 Nov 2019 · 0 repositories · arXiv:1911.12391
-
Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection 27 Nov 2019 · 0 repositories · arXiv:1911.11951
-
Autoencoding Undirected Molecular Graphs With Neural Networks 26 Nov 2019 · 1 repository · arXiv:2001.03517
-
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs 26 Nov 2019 · 1 repository · arXiv:1911.11390
-
Low Rank Factorization for Compact Multi-Head Self-Attention 26 Nov 2019 · 1 repository · arXiv:1912.00835
-
Password-conditioned Anonymization and Deanonymization with Face Identity Transformers 26 Nov 2019 · 1 repository · arXiv:1911.11759
-
Relevance-Promoting Language Model for Short-Text Conversation 26 Nov 2019 · 0 repositories · arXiv:1911.11489
-
Learning to Reuse Translations: Guiding Neural Machine Translation with Examples 25 Nov 2019 · 0 repositories · arXiv:1911.10732
-
Who did They Respond to? Conversation Structure Modeling using Masked Hierarchical Transformer 25 Nov 2019 · 1 repository · arXiv:1911.10666
-
Factorized Multimodal Transformer for Multimodal Sequential Learning 22 Nov 2019 · 0 repositories · arXiv:1911.09826
-
Improving N-gram Language Models with Pre-trained Deep Transformer 22 Nov 2019 · 0 repositories · arXiv:1911.10235
-
Neuron Interaction Based Representation Composition for Neural Machine Translation 22 Nov 2019 · 0 repositories · arXiv:1911.09877
-
Spectral Graph Transformer Networks for Brain Surface Parcellation 22 Nov 2019 · 0 repositories · arXiv:1911.10118
-
Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural Networks 21 Nov 2019 · 16 repositories · arXiv:1911.09737Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
WildMix Dataset and Spectro-Temporal Transformer Model for Monoaural Audio Source Separation 21 Nov 2019 · 0 repositories · arXiv:1911.09783
-
MarioNETte: Few-shot Face Reenactment Preserving Identity of Unseen Targets 19 Nov 2019 · 0 repositories · arXiv:1911.08139
-
Graph Transformer for Graph-to-Sequence Learning 18 Nov 2019 · 1 repository · arXiv:1911.07470Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Crowd Counting via Segmentation Guided Attention Networks and Curriculum Loss 18 Nov 2019 · 1 repository · arXiv:1911.07990
-
MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning 17 Nov 2019 · 3 repositories · arXiv:1911.09483Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Music theme recognition using CNN and self-attention 16 Nov 2019 · 0 repositories · arXiv:1911.07041
-
Evaluating robustness of language models for chief complaint extraction from patient-generated text 15 Nov 2019 · 0 repositories · arXiv:1911.06915
-
Interpreting chest X-rays via CNNs that exploit hierarchical disease dependencies and uncertainty labels 15 Nov 2019 · 2 repositories · arXiv:1911.06475
-
Selection-based Question Answering of an MOOC 15 Nov 2019 · 1 repository · arXiv:1911.07629
-
Sequential Recommendation with Relation-Aware Kernelized Self-Attention 15 Nov 2019 · 0 repositories · arXiv:1911.06478
-
Attention on Abstract Visual Reasoning 14 Nov 2019 · 0 repositories · arXiv:1911.05990
-
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA 14 Nov 2019 · 1 repository · arXiv:1911.06258
-
Character-based NMT with Transformer 12 Nov 2019 · 0 repositories · arXiv:1911.04997
-
SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery 12 Nov 2019 · 1 repository · arXiv:1911.04738Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Attending to Entities for Better Text Understanding 11 Nov 2019 · 0 repositories · arXiv:1911.04361
-
BP-Transformer: Modelling Long-Range Context via Binary Partitioning 11 Nov 2019 · 2 repositories · arXiv:1911.04070Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Disentangle, align and fuse for multimodal and semi-supervised image segmentation 11 Nov 2019 · 2 repositories · arXiv:1911.04417
-
Long-span language modeling for speech recognition 11 Nov 2019 · 0 repositories · arXiv:1911.04571
-
TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection 11 Nov 2019 · 2 repositories · arXiv:1911.04118
-
Distilling Knowledge Learned in BERT for Text Generation 10 Nov 2019 · 2 repositories · arXiv:1911.03829Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Learning to Few-Shot Learn Across Diverse Natural Language Classification Tasks 10 Nov 2019 · 2 repositories · arXiv:1911.03863
-
Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition 10 Nov 2019 · 0 repositories · arXiv:1911.04908