Methods › General › Attention Mechanisms › Attention › Papers, page 305
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 305 of 316: papers 30,401 to 30,500 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue 4 Dec 2019 · 0 repositories · arXiv:1912.10160
-
An Exploration of Data Augmentation and Sampling Techniques for Domain-Agnostic Question Answering 4 Dec 2019 · 0 repositories · arXiv:1912.02145
-
Enhancing Relation Extraction Using Syntactic Indicators and Sentential Contexts 4 Dec 2019 · 1 repository · arXiv:1912.01858
-
A Comparative Study of Pretrained Language Models on Thai Social Text Categorization 3 Dec 2019 · 0 repositories · arXiv:1912.01580
-
TU Wien @ TREC Deep Learning '19 -- Simple Contextualization for Re-ranking 3 Dec 2019 · 1 repository · arXiv:1912.01385
-
Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events 2 Dec 2019 · 0 repositories · arXiv:1912.02615
-
BERT for Large-scale Video Segment Classification with Test-time Augmentation 2 Dec 2019 · 0 repositories · arXiv:1912.01127
-
BLiMP: The Benchmark of Linguistic Minimal Pairs for English 2 Dec 2019 · 4 repositories · arXiv:1912.00582Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Leveraging Contextual Embeddings for Detecting Diachronic Semantic Shift 2 Dec 2019 · 0 repositories · arXiv:1912.01072
-
Long Distance Relationships without Time Travel: Boosting the Performance of a Sparse Predictive Autoencoder in Sequence Modeling 2 Dec 2019 · 1 repository · arXiv:1912.01116
-
Multi-Scale Self-Attention for Text Classification 2 Dec 2019 · 0 repositories · arXiv:1912.00544
-
Neural Academic Paper Generation 2 Dec 2019 · 1 repository · arXiv:1912.01982
-
Solving Arithmetic Word Problems Automatically Using Transformer and Unambiguous Representations 2 Dec 2019 · 1 repository · arXiv:1912.00871
-
Fast and Accurate Stochastic Gradient Estimation 1 Dec 2019 · 1 repository
-
Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks 1 Dec 2019 · 0 repositories
-
Perceiving the arrow of time in autoregressive motion 1 Dec 2019 · 0 repositories
-
Inducing Relational Knowledge from BERT 28 Nov 2019 · 0 repositories · arXiv:1911.12753
-
Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition 28 Nov 2019 · 0 repositories · arXiv:1911.12487
-
Automatic Generation of Headlines for Online Math Questions 27 Nov 2019 · 1 repository · arXiv:1912.00839
-
DeFINE: DEep Factorized INput Token Embeddings for Neural Sequence Modeling 27 Nov 2019 · 1 repository · arXiv:1911.12385
-
Do Attention Heads in BERT Track Syntactic Dependencies? 27 Nov 2019 · 1 repository · arXiv:1911.12246
-
Evaluating Commonsense in Pre-trained Language Models 27 Nov 2019 · 1 repository · arXiv:1911.11931
-
SimpleBooks: Long-term dependency book dataset with simplified English vocabulary for word-level language modeling 27 Nov 2019 · 0 repositories · arXiv:1911.12391
-
Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection 27 Nov 2019 · 0 repositories · arXiv:1911.11951
-
Autoencoding Undirected Molecular Graphs With Neural Networks 26 Nov 2019 · 1 repository · arXiv:2001.03517
-
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs 26 Nov 2019 · 1 repository · arXiv:1911.11390
-
Low Rank Factorization for Compact Multi-Head Self-Attention 26 Nov 2019 · 1 repository · arXiv:1912.00835
-
Password-conditioned Anonymization and Deanonymization with Face Identity Transformers 26 Nov 2019 · 1 repository · arXiv:1911.11759
-
Relevance-Promoting Language Model for Short-Text Conversation 26 Nov 2019 · 0 repositories · arXiv:1911.11489
-
Single Headed Attention RNN: Stop Thinking With Your Head 26 Nov 2019 · 5 repositories · arXiv:1911.11423
-
Learning to Reuse Translations: Guiding Neural Machine Translation with Examples 25 Nov 2019 · 0 repositories · arXiv:1911.10732
-
Who did They Respond to? Conversation Structure Modeling using Masked Hierarchical Transformer 25 Nov 2019 · 1 repository · arXiv:1911.10666
-
Factorized Multimodal Transformer for Multimodal Sequential Learning 22 Nov 2019 · 0 repositories · arXiv:1911.09826
-
Improving N-gram Language Models with Pre-trained Deep Transformer 22 Nov 2019 · 0 repositories · arXiv:1911.10235
-
Neuron Interaction Based Representation Composition for Neural Machine Translation 22 Nov 2019 · 0 repositories · arXiv:1911.09877
-
Spectral Graph Transformer Networks for Brain Surface Parcellation 22 Nov 2019 · 0 repositories · arXiv:1911.10118
-
Automatically Neutralizing Subjective Bias in Text 21 Nov 2019 · 1 repository · arXiv:1911.09709Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
Chemical-protein Interaction Extraction via Gaussian Probability Distribution and External Biomedical Knowledge 21 Nov 2019 · 1 repository · arXiv:1911.09487
-
Paraphrasing with Large Language Models 21 Nov 2019 · 0 repositories · arXiv:1911.09661
-
WildMix Dataset and Spectro-Temporal Transformer Model for Monoaural Audio Source Separation 21 Nov 2019 · 0 repositories · arXiv:1911.09783
-
Joint Emotion Label Space Modelling for Affect Lexica 20 Nov 2019 · 0 repositories · arXiv:1911.08782
-
MarioNETte: Few-shot Face Reenactment Preserving Identity of Unseen Targets 19 Nov 2019 · 0 repositories · arXiv:1911.08139
-
Towards Lingua Franca Named Entity Recognition with BERT 19 Nov 2019 · 0 repositories · arXiv:1912.01389
-
Towards non-toxic landscapes: Automatic toxic comment detection using DNN 19 Nov 2019 · 0 repositories · arXiv:1911.08395
-
Unsupervised Natural Question Answering with a Small Model 19 Nov 2019 · 0 repositories · arXiv:1911.08340
-
Graph Transformer for Graph-to-Sequence Learning 18 Nov 2019 · 1 repository · arXiv:1911.07470Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Improving Relation Classification by Entity Pair Graph 17 Nov 2019 · 0 repositories
-
MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning 17 Nov 2019 · 3 repositories · arXiv:1911.09483Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Unsupervised Visual Representation Learning with Increasing Object Shape Bias 17 Nov 2019 · 0 repositories · arXiv:1911.07272
-
Music theme recognition using CNN and self-attention 16 Nov 2019 · 0 repositories · arXiv:1911.07041
-
Robust Reading Comprehension with Linguistic Constraints via Posterior Regularization 16 Nov 2019 · 0 repositories · arXiv:1911.06948
-
Evaluating robustness of language models for chief complaint extraction from patient-generated text 15 Nov 2019 · 0 repositories · arXiv:1911.06915
-
Multi-attention Networks for Temporal Localization of Video-level Labels 15 Nov 2019 · 1 repository · arXiv:1911.06866
-
Selection-based Question Answering of an MOOC 15 Nov 2019 · 1 repository · arXiv:1911.07629
-
Sequential Recommendation with Relation-Aware Kernelized Self-Attention 15 Nov 2019 · 0 repositories · arXiv:1911.06478
-
Attention on Abstract Visual Reasoning 14 Nov 2019 · 0 repositories · arXiv:1911.05990
-
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA 14 Nov 2019 · 1 repository · arXiv:1911.06258
-
Adapting and evaluating a deep learning language model for clinical why-question answering 13 Nov 2019 · 0 repositories · arXiv:1911.05604
-
Compressive Transformers for Long-Range Sequence Modelling 13 Nov 2019 · 6 repositories · arXiv:1911.05507Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
Unsupervised Domain Adaptation on Reading Comprehension 13 Nov 2019 · 1 repository · arXiv:1911.06137
-
What do you mean, BERT? Assessing BERT as a Distributional Semantics Model 13 Nov 2019 · 0 repositories · arXiv:1911.05758
-
A Syntax-aware Multi-task Learning Framework for Chinese Semantic Role Labeling 12 Nov 2019 · 1 repository · arXiv:1911.04641
-
Character-based NMT with Transformer 12 Nov 2019 · 0 repositories · arXiv:1911.04997
-
SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery 12 Nov 2019 · 1 repository · arXiv:1911.04738Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Attending to Entities for Better Text Understanding 11 Nov 2019 · 0 repositories · arXiv:1911.04361
-
BP-Transformer: Modelling Long-Range Context via Binary Partitioning 11 Nov 2019 · 2 repositories · arXiv:1911.04070Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Disentangle, align and fuse for multimodal and semi-supervised image segmentation 11 Nov 2019 · 2 repositories · arXiv:1911.04417
-
Long-span language modeling for speech recognition 11 Nov 2019 · 0 repositories · arXiv:1911.04571
-
Meta Answering for Machine Reading 11 Nov 2019 · 0 repositories · arXiv:1911.04156
-
NegBERT: A Transfer Learning Approach for Negation Detection and Scope Resolution 11 Nov 2019 · 1 repository · arXiv:1911.04211
-
TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection 11 Nov 2019 · 2 repositories · arXiv:1911.04118
-
Understanding BERT performance in propaganda analysis 11 Nov 2019 · 0 repositories · arXiv:1911.04525
-
Distilling Knowledge Learned in BERT for Text Generation 10 Nov 2019 · 2 repositories · arXiv:1911.03829Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Effectiveness of self-supervised pre-training for speech recognition 10 Nov 2019 · 2 repositories · arXiv:1911.03912
-
Improving BERT Fine-tuning with Embedding Normalization 10 Nov 2019 · 0 repositories · arXiv:1911.03918
-
Improving Transformer Models by Reordering their Sublayers 10 Nov 2019 · 2 repositories · arXiv:1911.03864
-
INSET: Sentence Infilling with INter-SEntential Transformer 10 Nov 2019 · 1 repository · arXiv:1911.03892
-
Learning to Few-Shot Learn Across Diverse Natural Language Classification Tasks 10 Nov 2019 · 2 repositories · arXiv:1911.03863
-
Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition 10 Nov 2019 · 0 repositories · arXiv:1911.04908
-
RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers 10 Nov 2019 · 4 repositories · arXiv:1911.04942Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
Increasing Robustness to Spurious Correlations using Forgettable Examples 10 Nov 2019 · 0 repositories · arXiv:1911.03861
-
Syntax-Infused Transformer and BERT models for Machine Translation and Natural Language Understanding 10 Nov 2019 · 0 repositories · arXiv:1911.06156
-
TENER: Adapting Transformer Encoder for Named Entity Recognition 10 Nov 2019 · 6 repositories · arXiv:1911.04474Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Two-Headed Monster And Crossed Co-Attention Networks 10 Nov 2019 · 0 repositories · arXiv:1911.03897
-
Understanding Multi-Head Attention in Abstractive Summarization 10 Nov 2019 · 0 repositories · arXiv:1911.03898
-
Contextualized End-to-End Neural Entity Linking 10 Nov 2019 · 0 repositories · arXiv:1911.03834
-
Scalable Zero-shot Entity Linking with Dense Entity Retrieval 10 Nov 2019 · 3 repositories · arXiv:1911.03814
-
A Reinforced Generation of Adversarial Examples for Neural Machine Translation 9 Nov 2019 · 1 repository · arXiv:1911.03677
-
MKD: a Multi-Task Knowledge Distillation Approach for Pretrained Language Models 9 Nov 2019 · 0 repositories · arXiv:1911.03588
-
E-BERT: Efficient-Yet-Effective Entity Embeddings for BERT 9 Nov 2019 · 1 repository · arXiv:1911.03681
-
ConveRT: Efficient and Accurate Conversational Representations from Transformers 9 Nov 2019 · 5 repositories · arXiv:1911.03688Syntology 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
Multi-Perspective Inferrer: Reasoning Sentences Relationship from Holistic Perspective 9 Nov 2019 · 0 repositories · arXiv:1911.03668
-
The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational Agents 9 Nov 2019 · 0 repositories · arXiv:1911.03768
-
Zero-Shot Paraphrase Generation with Multilingual Language Models 9 Nov 2019 · 0 repositories · arXiv:1911.03597
-
Cross-Lingual Relevance Transfer for Document Retrieval 8 Nov 2019 · 0 repositories · arXiv:1911.02989
-
Graph-to-Graph Transformer for Transition-based Dependency Parsing 8 Nov 2019 · 1 repository · arXiv:1911.03561
-
How Language-Neutral is Multilingual BERT? 8 Nov 2019 · 1 repository · arXiv:1911.03310
-
Pretrained Language Models for Document-Level Neural Machine Translation 8 Nov 2019 · 0 repositories · arXiv:1911.03110
-
Question Generation from Paragraphs: A Tale of Two Hierarchical Models 8 Nov 2019 · 0 repositories · arXiv:1911.03407
-
Resurrecting Submodularity for Neural Text Generation 8 Nov 2019 · 0 repositories · arXiv:1911.03014