Methods › General › Attention Mechanisms › Attention › Papers, page 288
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 288 of 316: papers 28,701 to 28,800 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Towards Zero-Shot Conditional Summarization with Adaptive Multi-Task Fine-Tuning 1 Nov 2020 · 1 repository
-
Transformer-based Multi-Aspect Modeling for Multi-Aspect Multi-Sentiment Analysis 1 Nov 2020 · 0 repositories · arXiv:2011.00476
-
Visually-Grounded Planning without Vision: Language Models Infer Detailed Plans from High-level Instructions 1 Nov 2020 · 1 repository
-
Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution 31 Oct 2020 · 1 repository · arXiv:2011.00245
-
Neural Coreference Resolution for Arabic 31 Oct 2020 · 1 repository · arXiv:2011.00286
-
Understanding Pre-trained BERT for Aspect-based Sentiment Analysis 31 Oct 2020 · 2 repositories · arXiv:2011.00169
-
A Sui Generis QA Approach using RoBERTa for Adverse Drug Event Identification 30 Oct 2020 · 0 repositories · arXiv:2011.00057
-
Generating Radiology Reports via Memory-driven Transformer 30 Oct 2020 · 2 repositories · arXiv:2010.16056Syntology official (archive's flag): 1 ran · 26 ran (of which 13 constructed an object rather than computing a result; 20 with no instrument failure: 3 honoured, 3 violated, 14 with no contract checked; 6 where Syntology's instrument failed) · 5 unverified (of 31 harvested samples) · 25 pointer-only (licence)
-
Improving Dialogue Breakdown Detection with Semi-Supervised Learning 30 Oct 2020 · 0 repositories · arXiv:2011.00136
-
Semantic Labeling Using a Deep Contextualized Language Model 30 Oct 2020 · 1 repository · arXiv:2010.16037
-
SLM: Learning a Discourse Language Representation with Sentence Unshuffling 30 Oct 2020 · 0 repositories · arXiv:2010.16249
-
Target Word Masking for Location Metonymy Resolution 30 Oct 2020 · 1 repository · arXiv:2010.16097
-
Topic-Preserving Synthetic News Generation: An Adversarial Deep Reinforcement Learning Approach 30 Oct 2020 · 0 repositories · arXiv:2010.16324
-
VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation 30 Oct 2020 · 1 repository · arXiv:2010.16046
-
Combining Self-Training and Self-Supervised Learning for Unsupervised Disfluency Detection 29 Oct 2020 · 1 repository · arXiv:2010.15360
-
Contextual BERT: Conditioning the Language Model Using a Global State 29 Oct 2020 · 0 repositories · arXiv:2010.15778
-
Memory Attentive Fusion: External Language Model Integration for Transformer-based Sequence-to-Sequence Model 29 Oct 2020 · 0 repositories · arXiv:2010.15437
-
T-vectors: Weakly Supervised Speaker Identification Using Hierarchical Transformer Model 29 Oct 2020 · 0 repositories · arXiv:2010.16071
-
Tilde at WMT 2020: News Task Systems 29 Oct 2020 · 0 repositories · arXiv:2010.15423
-
Bayesian Methods for Semi-supervised Text Annotation 28 Oct 2020 · 0 repositories · arXiv:2010.14872
-
Detecting Stance in Media on Global Warming 28 Oct 2020 · 1 repository · arXiv:2010.15149
-
Fusion Models for Improved Visual Captioning 28 Oct 2020 · 0 repositories · arXiv:2010.15251
-
Learning to Unknot 28 Oct 2020 · 0 repositories · arXiv:2010.16263
-
The Volctrans Machine Translation System for WMT20 28 Oct 2020 · 0 repositories · arXiv:2010.14806
-
A Clarifying Question Selection System from NTES_ALONG in Convai3 Challenge 27 Oct 2020 · 0 repositories · arXiv:2010.14202
-
Fast Interleaved Bidirectional Sequence Generation 27 Oct 2020 · 1 repository · arXiv:2010.14481
-
FragmentVC: Any-to-Any Voice Conversion by End-to-End Extracting and Fusing Fine-Grained Voice Fragments With Attention 27 Oct 2020 · 2 repositories · arXiv:2010.14150Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering 27 Oct 2020 · 1 repository · arXiv:2010.14095
-
Multimodal Emotion Recognition with Transformer-Based Self Supervised Feature Fusion 27 Oct 2020 · 1 repository
-
Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discriminators 27 Oct 2020 · 0 repositories · arXiv:2010.14151
-
To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging 27 Oct 2020 · 0 repositories · arXiv:2010.14042
-
Unmasking Contextual Stereotypes: Measuring and Mitigating BERT's Gender Bias 27 Oct 2020 · 1 repository · arXiv:2010.14534
-
Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping 26 Oct 2020 · 1 repository · arXiv:2010.13369
-
Controlled Molecule Generator for Optimizing Multiple Chemical Properties 26 Oct 2020 · 1 repository · arXiv:2010.13908Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
FastFormers: Highly Efficient Transformer Models for Natural Language Understanding 26 Oct 2020 · 2 repositories · arXiv:2010.13382Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Fine-grained Information Status Classification Using Discourse Context-Aware BERT 26 Oct 2020 · 1 repository · arXiv:2010.14759
-
Graph Transformer Networks with Syntactic and Semantic Structures for Event Argument Extraction 26 Oct 2020 · 0 repositories · arXiv:2010.13391
-
Peak Detection On Data Independent Acquisition Mass Spectrometry Data With Semisupervised Convolutional Transformers 26 Oct 2020 · 0 repositories · arXiv:2010.13841
-
Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining 26 Oct 2020 · 1 repository · arXiv:2010.13826
-
UPB at SemEval-2020 Task 12: Multilingual Offensive Language Detection on Social Media by Fine-tuning a Variety of BERT-based Models 26 Oct 2020 · 0 repositories · arXiv:2010.13609
-
Attention is All You Need in Speech Separation 25 Oct 2020 · 4 repositories · arXiv:2010.13154
-
Commonsense knowledge adversarial dataset that challenges ELECTRA 25 Oct 2020 · 0 repositories · arXiv:2010.13049
-
Contextualized Word Embeddings Encode Aspects of Human-Like Word Sense Knowledge 25 Oct 2020 · 0 repositories · arXiv:2010.13057
-
CRAB: Class Representation Attentive BERT for Hate Speech Identification in Social Media 25 Oct 2020 · 0 repositories · arXiv:2010.13028
-
Two-stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding 25 Oct 2020 · 1 repository · arXiv:2010.13105
-
Char2Subword: Extending the Subword Embedding Space Using Robust Character Compositionality 24 Oct 2020 · 0 repositories · arXiv:2010.12730
-
COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval 24 Oct 2020 · 1 repository · arXiv:2010.12800
-
Effective Distant Supervision for Temporal Relation Extraction 24 Oct 2020 · 2 repositories · arXiv:2010.12755
-
Hierarchical Transformer for Task Oriented Dialog Systems 24 Oct 2020 · 2 repositories · arXiv:2011.08067
-
Measuring Association Between Labels and Free-Text Rationales 24 Oct 2020 · 1 repository · arXiv:2010.12762Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Jointly Optimizing State Operation Prediction and Value Generation for Dialogue State Tracking 24 Oct 2020 · 2 repositories · arXiv:2010.14061
-
Open-Domain Dialogue Generation Based on Pre-trained Language Models 24 Oct 2020 · 0 repositories · arXiv:2010.12780
-
Pre-trained Summarization Distillation 24 Oct 2020 · 1 repository · arXiv:2010.13002
-
Rethinking embedding coupling in pre-trained language models 24 Oct 2020 · 4 repositories · arXiv:2010.12821
-
Unsupervised Paraphrasing with Pretrained Language Models 24 Oct 2020 · 0 repositories · arXiv:2010.12885
-
Pre-training Graph Transformer with Multimodal Side Information for Recommendation 23 Oct 2020 · 0 repositories · arXiv:2010.12284
-
A Simple Approach for Handling Out-of-Vocabulary Identifiers in Deep Learning for Source Code 23 Oct 2020 · 1 repository · arXiv:2010.12663
-
BARThez: a Skilled Pretrained French Sequence-to-Sequence Model 23 Oct 2020 · 5 repositories · arXiv:2010.12321
-
Deep Learning Framework for Measuring the Digital Strategy of Companies from Earnings Calls 23 Oct 2020 · 1 repository · arXiv:2010.12418
-
Did You Ask a Good Question? A Cross-Domain Question Intention Classification Benchmark for Text-to-SQL 23 Oct 2020 · 1 repository · arXiv:2010.12634
-
Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer 23 Oct 2020 · 1 repository · arXiv:2010.12180
-
ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding 23 Oct 2020 · 2 repositories · arXiv:2010.12148
-
GiBERT: Introducing Linguistic Knowledge into BERT through a Lightweight Gated Injection Method 23 Oct 2020 · 0 repositories · arXiv:2010.12532
-
GraphSpeech: Syntax-Aware Graph Attention Network For Neural Speech Synthesis 23 Oct 2020 · 0 repositories · arXiv:2010.12423
-
HateBERT: Retraining BERT for Abusive Language Detection in English 23 Oct 2020 · 1 repository · arXiv:2010.12472
-
LightSeq: A High Performance Inference Library for Transformers 23 Oct 2020 · 1 repository · arXiv:2010.13887
-
Long Document Ranking with Query-Directed Sparse Transformer 23 Oct 2020 · 1 repository · arXiv:2010.12683
-
Multilingual BERT Post-Pretraining Alignment 23 Oct 2020 · 0 repositories · arXiv:2010.12547
-
On the Transformer Growth for Progressive BERT Training 23 Oct 2020 · 0 repositories · arXiv:2010.12562
-
Posterior Differential Regularization with f-divergence for Improving Model Robustness 23 Oct 2020 · 2 repositories · arXiv:2010.12638
-
Pre-training with Meta Learning for Chinese Word Segmentation 23 Oct 2020 · 0 repositories · arXiv:2010.12272
-
ST-BERT: Cross-modal Language Model Pre-training For End-to-end Spoken Language Understanding 23 Oct 2020 · 0 repositories · arXiv:2010.12283
-
Stabilizing Transformer-Based Action Sequence Generation For Q-Learning 23 Oct 2020 · 0 repositories · arXiv:2010.12698
-
Topic Modeling with Contextualized Word Representation Clusters 23 Oct 2020 · 0 repositories · arXiv:2010.12626
-
Abstracting the Traffic of Nonlinear Event-Triggered Control Systems 23 Oct 2020 · 0 repositories · arXiv:2010.12341
-
Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention 23 Oct 2020 · 1 repository · arXiv:2010.12155
-
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale 22 Oct 2020 · 158 repositories · arXiv:2010.11929Syntology official: harvested, nothing ran · 307 ran (of which 165 constructed an object rather than computing a result; 286 with no instrument failure: 8 honoured, 2 violated, 276 with no contract checked; 21 where Syntology's instrument failed) · 112 unverified (of 419 harvested samples) · 154 pointer-only (licence)
-
Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset 22 Oct 2020 · 0 repositories · arXiv:2010.11395
-
Distilling Dense Representations for Ranking using Tightly-Coupled Teachers 22 Oct 2020 · 2 repositories · arXiv:2010.11386
-
How Phonotactics Affect Multilingual and Zero-shot ASR Performance 22 Oct 2020 · 1 repository · arXiv:2010.12104
-
Improving BERT Performance for Aspect-Based Sentiment Analysis 22 Oct 2020 · 2 repositories · arXiv:2010.11731Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 2 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Exploiting News Article Structure for Automatic Corpus Generation of Entailment Datasets 22 Oct 2020 · 1 repository · arXiv:2010.11574
-
Knowledge Distillation for BERT Unsupervised Domain Adaptation 22 Oct 2020 · 1 repository · arXiv:2010.11478
-
Language Models are Open Knowledge Graphs 22 Oct 2020 · 2 repositories · arXiv:2010.11967Syntology 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
mT5: A massively multilingual pre-trained text-to-text transformer 22 Oct 2020 · 8 repositories · arXiv:2010.11934Syntology official (archive's flag): 5 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations 22 Oct 2020 · 1 repository · arXiv:2010.11358
-
Scientific Claim Verification with VERT5ERINI 22 Oct 2020 · 0 repositories · arXiv:2010.11930
-
Self-Alignment Pretraining for Biomedical Entity Representations 22 Oct 2020 · 1 repository · arXiv:2010.11784
-
Towards Fully Bilingual Deep Language Modeling 22 Oct 2020 · 0 repositories · arXiv:2010.11639
-
UniCase -- Rethinking Casing in Language Models 22 Oct 2020 · 0 repositories · arXiv:2010.11936
-
Analyzing the Source and Target Contributions to Predictions in Neural Machine Translation 21 Oct 2020 · 1 repository · arXiv:2010.10907
-
Detection of COVID-19 informative tweets using RoBERTa 21 Oct 2020 · 0 repositories · arXiv:2010.11238
-
A Simple and Efficient Multi-Task Learning Approach for Conditioned Dialogue Generation 21 Oct 2020 · 1 repository · arXiv:2010.11140
-
German's Next Language Model 21 Oct 2020 · 1 repository · arXiv:2010.10906
-
Latte-Mix: Measuring Sentence Semantic Similarity with Latent Categorical Mixtures 21 Oct 2020 · 0 repositories · arXiv:2010.11351
-
Multi-Domain Dialogue State Tracking based on State Graph 21 Oct 2020 · 0 repositories · arXiv:2010.11137
-
Multi-Unit Transformers for Neural Machine Translation 21 Oct 2020 · 1 repository · arXiv:2010.10743
-
TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog 21 Oct 2020 · 0 repositories · arXiv:2010.10839
-
Token Drop mechanism for Neural Machine Translation 21 Oct 2020 · 1 repository · arXiv:2010.11018
-
Transferable Graph Optimizers for ML Compilers 21 Oct 2020 · 0 repositories · arXiv:2010.12438