Methods › General › Attention Modules › Multi-Head Attention › Papers, page 240
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 240 of 249: papers 23,901 to 24,000 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Unsupervised Labeled Parsing with Deep Inside-Outside Recursive Autoencoders 1 Nov 2019 · 0 repositories
-
What Does This Word Mean? Explaining Contextualized Embeddings with Natural Language Definition 1 Nov 2019 · 0 repositories
-
When Choosing Plausible Alternatives, Clever Hans can be Clever 1 Nov 2019 · 0 repositories · arXiv:1911.00225
-
Attention Is All You Need for Chinese Word Segmentation 31 Oct 2019 · 1 repository · arXiv:1910.14537
-
DiaNet: BERT and Hierarchical Attention Multi-Task Learning of Fine-Grained Dialect 31 Oct 2019 · 0 repositories · arXiv:1910.14243
-
Do Multi-hop Readers Dream of Reasoning Chains? 31 Oct 2019 · 1 repository · arXiv:1910.14520Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Document-level Neural Machine Translation with Associated Memory Network 31 Oct 2019 · 0 repositories · arXiv:1910.14528
-
Human-centric Metric for Accelerating Pathology Reports Annotation 31 Oct 2019 · 0 repositories · arXiv:1911.01226
-
Image-Conditioned Graph Generation for Road Network Extraction 31 Oct 2019 · 3 repositories · arXiv:1910.14388
-
LIMIT-BERT : Linguistic Informed Multi-Task BERT 31 Oct 2019 · 0 repositories · arXiv:1910.14296
-
Multi-Stage Document Ranking with BERT 31 Oct 2019 · 3 repositories · arXiv:1910.14424Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
NAT: Neural Architecture Transformer for Accurate and Compact Architectures 31 Oct 2019 · 1 repository · arXiv:1910.14488
-
Neural Assistant: Joint Action Prediction, Response Generation, and Latent Knowledge Reasoning 31 Oct 2019 · 1 repository · arXiv:1910.14613
-
Parameter Sharing Decoder Pair for Auto Composing 31 Oct 2019 · 0 repositories · arXiv:1910.14270
-
Positional Attention-based Frame Identification with BERT: A Deep Learning Approach to Target Disambiguation and Semantic Frame Selection 31 Oct 2019 · 0 repositories · arXiv:1910.14549
-
Masked Language Model Scoring 31 Oct 2019 · 6 repositories · arXiv:1910.14659Syntology official: no sample here; runs from other or unrecorded repositories · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 2 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 3 pointer-only (licence)
-
Transfer Learning from Transformers to Fake News Challenge Stance Detection (FNC-1) Task 31 Oct 2019 · 0 repositories · arXiv:1910.14353
-
An Augmented Transformer Architecture for Natural Language Generation Tasks 30 Oct 2019 · 0 repositories · arXiv:1910.13634
-
Discourse-Aware Neural Extractive Text Summarization 30 Oct 2019 · 1 repository · arXiv:1910.14142Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Lightweight and Efficient End-to-End Speech Recognition Using Low-Rank Transformer 30 Oct 2019 · 0 repositories · arXiv:1910.13923
-
Lsh-sampling Breaks the Computation Chicken-and-egg Loop in Adaptive Stochastic Gradient Estimation 30 Oct 2019 · 0 repositories · arXiv:1910.14162
-
Phenotyping of Clinical Notes with Improved Document Classification Models Using Contextualized Neural Language Models 30 Oct 2019 · 2 repositories · arXiv:1910.13664
-
Time to Take Emoji Seriously: They Vastly Improve Casual Conversational Models 30 Oct 2019 · 0 repositories · arXiv:1910.13793
-
Transformer-based Cascaded Multimodal Speech Translation 29 Oct 2019 · 0 repositories · arXiv:1910.13215
-
Sentence Embeddings for Russian NLU 29 Oct 2019 · 1 repository · arXiv:1910.13291
-
An Empirical Study of Generation Order for Machine Translation 29 Oct 2019 · 0 repositories · arXiv:1910.13437
-
Big Bidirectional Insertion Representations for Documents 29 Oct 2019 · 0 repositories · arXiv:1910.13034
-
Inducing brain-relevant bias in natural language processing models 29 Oct 2019 · 1 repository · arXiv:1911.03268Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Learning Rich Image Region Representation for Visual Question Answering 29 Oct 2019 · 0 repositories · arXiv:1910.13077
-
A BERT-Based Transfer Learning Approach for Hate Speech Detection in Online Social Media 28 Oct 2019 · 2 repositories · arXiv:1910.12574Syntology 0 ran · 1 unverified (of 1 harvested sample)
-
A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systems 28 Oct 2019 · 0 repositories · arXiv:1910.12995
-
Fine-Grained Object Detection over Scientific Document Images with Region Embeddings 28 Oct 2019 · 0 repositories · arXiv:1910.12462
-
Modeling Inter-Speaker Relationship in XLNet for Contextual Spoken Language Understanding 28 Oct 2019 · 0 repositories · arXiv:1910.12531
-
Sequence-to-sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding 28 Oct 2019 · 1 repository · arXiv:1910.12740
-
Transformer-Transducer: End-to-End Speech Recognition with Self-Attention 28 Oct 2019 · 1 repository · arXiv:1910.12977
-
What does BERT Learn from Multiple-Choice Reading Comprehension Datasets? 28 Oct 2019 · 0 repositories · arXiv:1910.12391
-
An Adaptive and Momental Bound Method for Stochastic Learning 27 Oct 2019 · 2 repositories · arXiv:1910.12249Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Word-level Textual Adversarial Attacking as Combinatorial Optimization 27 Oct 2019 · 1 repository · arXiv:1910.12196Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Thieves on Sesame Street! Model Extraction of BERT-based APIs 27 Oct 2019 · 1 repository · arXiv:1910.12366
-
Training ASR models by Generation of Contextual Information 27 Oct 2019 · 0 repositories · arXiv:1910.12367
-
DENS: A Dataset for Multi-class Emotion Analysis 25 Oct 2019 · 0 repositories · arXiv:1910.11769
-
HUBERT Untangles BERT to Improve Transfer across NLP Tasks 25 Oct 2019 · 1 repository · arXiv:1910.12647
-
L2RS: A Learning-to-Rescore Mechanism for Automatic Speech Recognition 25 Oct 2019 · 0 repositories · arXiv:1910.11496
-
Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders 25 Oct 2019 · 7 repositories · arXiv:1910.12638Syntology community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
On the Cross-lingual Transferability of Monolingual Representations 25 Oct 2019 · 7 repositories · arXiv:1910.11856Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering 25 Oct 2019 · 0 repositories · arXiv:1910.11559
-
Towards Online End-to-end Transformer Automatic Speech Recognition 25 Oct 2019 · 0 repositories · arXiv:1910.11871
-
An Empirical Study of Efficient ASR Rescoring with Transformers 24 Oct 2019 · 0 repositories · arXiv:1910.11450
-
Combining Acoustics, Content and Interaction Features to Find Hot Spots in Meetings 24 Oct 2019 · 0 repositories · arXiv:1910.10869
-
ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit 24 Oct 2019 · 3 repositories · arXiv:1910.10909Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Promoting the Knowledge of Source Syntax in Transformer NMT Is Not Needed 24 Oct 2019 · 0 repositories · arXiv:1910.11218
-
A Transformer with Interleaved Self-attention and Convolution for Hybrid Acoustic Models 23 Oct 2019 · 1 repository · arXiv:1910.10352
-
Controlling the Output Length of Neural Machine Translation 23 Oct 2019 · 0 repositories · arXiv:1910.10408
-
Correction of Automatic Speech Recognition with Transformer Sequence-to-sequence Model 23 Oct 2019 · 0 repositories · arXiv:1910.10697
-
Deja-vu: Double Feature Presentation and Iterated Loss in Deep Transformer Networks 23 Oct 2019 · 2 repositories · arXiv:1910.10324
-
Emergent Properties of Finetuned Language Representation Models 23 Oct 2019 · 0 repositories · arXiv:1910.10832
-
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 23 Oct 2019 · 57 repositories · arXiv:1910.10683Syntology 21 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 1 honoured, 0 violated, 19 with no contract checked; 1 where Syntology's instrument failed) · 10 unverified (of 31 harvested samples)
-
Hierarchical Transformers for Long Document Classification 23 Oct 2019 · 3 repositories · arXiv:1910.10781
-
Relation Module for Non-answerable Prediction on Question Answering 23 Oct 2019 · 0 repositories · arXiv:1910.10843
-
Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks 23 Oct 2019 · 0 repositories · arXiv:1910.10387
-
TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation 23 Oct 2019 · 0 repositories · arXiv:1911.05186
-
Complex Transformer: A Framework for Modeling Complex-Valued Sequence 22 Oct 2019 · 1 repository · arXiv:1910.10202Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Exchangeable deep neural networks for set-to-set matching and learning 22 Oct 2019 · 2 repositories · arXiv:1910.09972
-
Depth-Adaptive Transformer 22 Oct 2019 · 0 repositories · arXiv:1910.10073
-
Improving Transformer-based Speech Recognition Using Unsupervised Pre-training 22 Oct 2019 · 1 repository · arXiv:1910.09932
-
MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension 22 Oct 2019 · 1 repository · arXiv:1910.09753Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Sequence-to-sequence Singing Synthesis Using the Feed-forward Transformer 22 Oct 2019 · 0 repositories · arXiv:1910.09989
-
Transformer-based Acoustic Modeling for Hybrid Speech Recognition 22 Oct 2019 · 0 repositories · arXiv:1910.09799
-
Learning to Make Generalizable and Diverse Predictions for Retrosynthesis 21 Oct 2019 · 0 repositories · arXiv:1910.09688
-
Transformer-CNN: Fast and Reliable tool for QSAR 21 Oct 2019 · 1 repository · arXiv:1911.06603
-
Personalized Graph Neural Networks with Attention Mechanism for Session-Aware Recommendation 20 Oct 2019 · 3 repositories · arXiv:1910.08887
-
Targeted Estimation of Heterogeneous Treatment Effect in Observational Survival Analysis 20 Oct 2019 · 1 repository · arXiv:1910.08877
-
MonaLog: a Lightweight System for Natural Language Inference Based on Monotonicity 19 Oct 2019 · 1 repository · arXiv:1910.08772
-
XL-Editor: Post-editing Sentences with XLNet 19 Oct 2019 · 0 repositories · arXiv:1910.10479
-
A Mutual Information Maximization Perspective of Language Representation Learning 18 Oct 2019 · 0 repositories · arXiv:1910.08350
-
Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System 18 Oct 2019 · 0 repositories · arXiv:1910.08381
-
BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge 17 Oct 2019 · 0 repositories · arXiv:1910.07713
-
Fully Quantized Transformer for Machine Translation 17 Oct 2019 · 0 repositories · arXiv:1910.10485
-
Measuring semantic similarity of clinical trial outcomes using deep pre-trained language representations 17 Oct 2019 · 0 repositories
-
Predicting retrosynthetic pathways using a combined linguistic model and hyper-graph exploration strategy 17 Oct 2019 · 0 repositories · arXiv:1910.08036
-
Question Classification with Deep Contextualized Transformer 17 Oct 2019 · 1 repository · arXiv:1910.10492
-
Universal Text Representation from BERT: An Empirical Study 17 Oct 2019 · 0 repositories · arXiv:1910.07973
-
BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model Performance 16 Oct 2019 · 1 repository · arXiv:1910.07181Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Bridging the Knowledge Gap: Enhancing Question Answering with World and Domain Knowledge 16 Oct 2019 · 0 repositories · arXiv:1910.07429
-
Efficiency through Auto-Sizing: Notre Dame NLP's Submission to the WNGT 2019 Efficiency Task 16 Oct 2019 · 0 repositories · arXiv:1910.07134
-
Evolution of transfer learning in natural language processing 16 Oct 2019 · 0 repositories · arXiv:1910.07370
-
Imperial College London Submission to VATEX Video Captioning Task 16 Oct 2019 · 0 repositories · arXiv:1910.07482
-
Injecting Hierarchy with U-Net Transformers 16 Oct 2019 · 2 repositories · arXiv:1910.10488
-
Analyzing the Forgetting Problem in the Pretrain-Finetuning of Dialogue Response Models 16 Oct 2019 · 0 repositories · arXiv:1910.07117
-
Transformer ASR with Contextual Block Processing 16 Oct 2019 · 0 repositories · arXiv:1910.07204
-
Using Whole Document Context in Neural Machine Translation 16 Oct 2019 · 0 repositories · arXiv:1910.07481
-
Aligning Cross-Lingual Entities with Multi-Aspect Information 15 Oct 2019 · 1 repository · arXiv:1910.06575
-
Answering Complex Open-domain Questions Through Iterative Query Generation 15 Oct 2019 · 1 repository · arXiv:1910.07000Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving 15 Oct 2019 · 3 repositories · arXiv:1910.06611Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples)
-
Facebook AI's WAT19 Myanmar-English Translation Task Submission 15 Oct 2019 · 0 repositories · arXiv:1910.06848
-
Structured Pruning of a BERT-based Question Answering Model 14 Oct 2019 · 0 repositories · arXiv:1910.06360
-
Q8BERT: Quantized 8Bit BERT 14 Oct 2019 · 5 repositories · arXiv:1910.06188Syntology official (archive's flag): 5 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
STANCY: Stance Classification Based on Consistency Cues 14 Oct 2019 · 1 repository · arXiv:1910.06048
-
Transformers without Tears: Improving the Normalization of Self-Attention 14 Oct 2019 · 5 repositories · arXiv:1910.05895Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Whatcha lookin' at? DeepLIFTing BERT's Attention in Question Answering 14 Oct 2019 · 1 repository · arXiv:1910.06431