Methods › General › Attention Modules › Multi-Head Attention › Papers, page 237
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 237 of 249: papers 23,601 to 23,700 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A BERT based Sentiment Analysis and Key Entity Detection Approach for Online Financial Texts 14 Jan 2020 · 0 repositories · arXiv:2001.05326
-
Auto Completion of User Interface Layout Design Using Transformer-Based Tree Decoders 14 Jan 2020 · 0 repositories · arXiv:2001.05308
-
The problems with using STNs to align CNN feature maps 14 Jan 2020 · 0 repositories · arXiv:2001.05858
-
AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search 13 Jan 2020 · 1 repository · arXiv:2001.04246Syntology 0 ran · 3 unverified (of 3 harvested samples)
-
Reformer: The Efficient Transformer 13 Jan 2020 · 10 repositories · arXiv:2001.04451Syntology community repositories only · 6 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Représentations lexicales pour la détection non supervisée d'événements dans un flux de tweets : étude sur des corpus français et anglais 13 Jan 2020 · 1 repository · arXiv:2001.04139
-
Urdu-English Machine Transliteration using Neural Networks 12 Jan 2020 · 0 repositories · arXiv:2001.05296
-
Exploring and Improving Robustness of Multi Task Deep Neural Networks via Domain Agnostic Defenses 11 Jan 2020 · 1 repository · arXiv:2001.05286
-
PatentTransformer-2: Controlling Patent Text Generation by Structural Metadata 11 Jan 2020 · 0 repositories · arXiv:2001.03708
-
Resolving the Scope of Speculation and Negation using Transformer-Based Architectures 9 Jan 2020 · 1 repository · arXiv:2001.02885
-
Spatial-Temporal Transformer Networks for Traffic Flow Forecasting 9 Jan 2020 · 1 repository · arXiv:2001.02908
-
Streaming automatic speech recognition with the transformer model 8 Jan 2020 · 0 repositories · arXiv:2001.02674
-
To Transfer or Not to Transfer: Misclassification Attacks Against Transfer Learned Text Classifiers 8 Jan 2020 · 0 repositories · arXiv:2001.02438
-
Knowledge-aware Attention Network for Protein-Protein Interaction Extraction 7 Jan 2020 · 1 repository · arXiv:2001.02091
-
RECAST: Interactive Auditing of Automatic Toxicity Detection Models 7 Jan 2020 · 0 repositories · arXiv:2001.01819
-
Improving Entity Linking by Modeling Latent Entity Type Information 6 Jan 2020 · 0 repositories · arXiv:2001.01447
-
FDFtNet: Facing Off Fake Images using Fake Detection Fine-tuning Network 5 Jan 2020 · 2 repositories · arXiv:2001.01265Syntology official (archive's flag): 4 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples)
-
Explain and Improve: LRP-Inference Fine-Tuning for Image Captioning Models 4 Jan 2020 · 1 repository · arXiv:2001.01037
-
Learning Accurate Integer Transformer Machine-Translation Models 3 Jan 2020 · 0 repositories · arXiv:2001.00926
-
Multi-Layer Content Interaction Through Quaternion Product For Visual Question Answering 3 Jan 2020 · 0 repositories · arXiv:2001.05840
-
Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation 3 Jan 2020 · 1 repository · arXiv:2001.00891
-
Representing Unordered Data Using Complex-Weighted Multiset Automata 2 Jan 2020 · 0 repositories · arXiv:2001.00610
-
BERT-AL: BERT for Arbitrarily Long Document Understanding 1 Jan 2020 · 0 repositories
-
NEURAL EXECUTION ENGINES 1 Jan 2020 · 0 repositories
-
Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring 1 Jan 2020 · 2 repositories
-
Stacked DeBERT: All Attention in Incomplete Data for Text Classification 1 Jan 2020 · 1 repository · arXiv:2001.00137Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Deep Attentive Ranking Networks for Learning to Order Sentences 31 Dec 2019 · 0 repositories · arXiv:2001.00056
-
EEG based Continuous Speech Recognition using Transformers 31 Dec 2019 · 0 repositories · arXiv:2001.00501
-
oLMpics -- On what Language Model Pre-training Captures 31 Dec 2019 · 2 repositories · arXiv:1912.13283
-
OTEANN: Estimating the Transparency of Orthographies with an Artificial Neural Network 31 Dec 2019 · 2 repositories · arXiv:1912.13321
-
AraNet: A Deep Learning Toolkit for Arabic Social Media 30 Dec 2019 · 1 repository · arXiv:1912.13072
-
AutoDiscern: Rating the Quality of Online Health Information with Hierarchical Encoder Attention-based Neural Networks 30 Dec 2019 · 1 repository · arXiv:1912.12999
-
All-in-One Image-Grounded Conversational Agents 28 Dec 2019 · 0 repositories · arXiv:1912.12394
-
Energy-based Graph Convolutional Networks for Scoring Protein Docking Models 28 Dec 2019 · 0 repositories · arXiv:1912.12476
-
Clinical XLNet: Modeling Sequential Clinical Notes and Predicting Prolonged Mechanical Ventilation 27 Dec 2019 · 3 repositories · arXiv:1912.11975
-
Encoding word order in complex embeddings 27 Dec 2019 · 1 repository · arXiv:1912.12333
-
Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention 27 Dec 2019 · 1 repository · arXiv:1912.11959
-
Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection 25 Dec 2019 · 2 repositories · arXiv:1912.11637Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 1 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Leveraging Lead Bias for Zero-shot Abstractive News Summarization 25 Dec 2019 · 0 repositories · arXiv:1912.11602
-
Improving Abstractive Text Summarization with History Aggregation 24 Dec 2019 · 0 repositories · arXiv:1912.11046
-
Multi-Graph Transformer for Free-Hand Sketch Recognition 24 Dec 2019 · 1 repository · arXiv:1912.11258Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
Probing the phonetic and phonological knowledge of tones in Mandarin TTS models 23 Dec 2019 · 1 repository · arXiv:1912.10915
-
end-to-end training of a large vocabulary end-to-end speech recognition system 22 Dec 2019 · 0 repositories · arXiv:1912.11040
-
Harnessing Evolution of Multi-Turn Conversations for Effective Answer Retrieval 22 Dec 2019 · 1 repository · arXiv:1912.10554
-
Learning and Evaluating Contextual Embedding of Source Code 21 Dec 2019 · 2 repositories · arXiv:2001.00059
-
Are Transformers universal approximators of sequence-to-sequence functions? 20 Dec 2019 · 0 repositories · arXiv:1912.10077
-
Axial Attention in Multidimensional Transformers 20 Dec 2019 · 3 repositories · arXiv:1912.12180Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
ET-USB: Transformer-Based Sequential Behavior Modeling for Inbound Customer Service 20 Dec 2019 · 0 repositories · arXiv:1912.10852
-
Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model 20 Dec 2019 · 0 repositories · arXiv:1912.09637
-
Shareable Representations for Search Query Understanding 20 Dec 2019 · 0 repositories · arXiv:2001.04345
-
BERTje: A Dutch BERT Model 19 Dec 2019 · 2 repositories · arXiv:1912.09582Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
CJRC: A Reliable Human-Annotated Benchmark DataSet for Chinese Judicial Reading Comprehension 19 Dec 2019 · 0 repositories · arXiv:1912.09156
-
Neural Simile Recognition with Cyclic Multitask Learning and Local Attention 19 Dec 2019 · 1 repository · arXiv:1912.09084
-
Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting 19 Dec 2019 · 36 repositories · arXiv:1912.09363Syntology 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
A Multi-task Learning Model for Chinese-oriented Aspect Polarity Classification and Aspect Term Extraction 17 Dec 2019 · 6 repositories · arXiv:1912.07976
-
Cross-Lingual Ability of Multilingual BERT: An Empirical Study 17 Dec 2019 · 0 repositories · arXiv:1912.07840
-
Meshed-Memory Transformer for Image Captioning 17 Dec 2019 · 2 repositories · arXiv:1912.08226Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
The performance evaluation of Multi-representation in the Deep Learning models for Relation Extraction Task 17 Dec 2019 · 0 repositories · arXiv:1912.08290
-
Learning Malware Representation based on Execution Sequences 16 Dec 2019 · 0 repositories · arXiv:1912.07250
-
Multilingual is not enough: BERT for Finnish 15 Dec 2019 · 1 repository · arXiv:1912.07076
-
Robust Named Entity Recognition with Truecasing Pretraining 15 Dec 2019 · 0 repositories · arXiv:1912.07095
-
BERTQA -- Attention on Steroids 14 Dec 2019 · 0 repositories · arXiv:1912.10435
-
Towards Robust Toxic Content Classification 14 Dec 2019 · 1 repository · arXiv:1912.06872
-
Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining 14 Dec 2019 · 2 repositories · arXiv:1912.06813
-
TopoAct: Visually Exploring the Shape of Activations in Deep Learning 13 Dec 2019 · 1 repository · arXiv:1912.06332
-
WaLDORf: Wasteless Language-model Distillation On Reading-comprehension 13 Dec 2019 · 0 repositories · arXiv:1912.06638
-
BERT has a Moral Compass: Improvements of ethical and moral values of machines 11 Dec 2019 · 0 repositories · arXiv:1912.05238
-
Encoding Musical Style with Transformer Autoencoders 10 Dec 2019 · 0 repositories · arXiv:1912.05537
-
Unsupervised Transfer Learning via BERT Neuron Selection 10 Dec 2019 · 0 repositories · arXiv:1912.05308
-
Learning a Layout Transfer Network for Context Aware Object Detection 9 Dec 2019 · 0 repositories · arXiv:1912.03865
-
Transformer Based Reinforcement Learning For Games 9 Dec 2019 · 0 repositories · arXiv:1912.03918
-
Bidirectional Scene Text Recognition with a Single Decoder 8 Dec 2019 · 1 repository · arXiv:1912.03656
-
Adversarial Analysis of Natural Language Inference Systems 7 Dec 2019 · 0 repositories · arXiv:1912.03441
-
Personalized Patent Claim Generation and Measurement 7 Dec 2019 · 0 repositories · arXiv:1912.03502
-
Semantic Mask for Transformer based End-to-End Speech Recognition 6 Dec 2019 · 1 repository · arXiv:1912.03010
-
Synchronous Transformers for End-to-End Speech Recognition 6 Dec 2019 · 0 repositories · arXiv:1912.02958
-
Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks 6 Dec 2019 · 0 repositories · arXiv:1912.03063
-
Why are Adaptive Methods Good for Attention Models? 6 Dec 2019 · 0 repositories · arXiv:1912.03194
-
Self-Supervised Contextual Language Representation of Radiology Reports to Improve the Identification of Communication Urgency 5 Dec 2019 · 0 repositories · arXiv:1912.02703
-
Acquiring Knowledge from Pre-trained Model to Neural Machine Translation 4 Dec 2019 · 0 repositories · arXiv:1912.01774
-
AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue 4 Dec 2019 · 0 repositories · arXiv:1912.10160
-
An Exploration of Data Augmentation and Sampling Techniques for Domain-Agnostic Question Answering 4 Dec 2019 · 0 repositories · arXiv:1912.02145
-
Enhancing Relation Extraction Using Syntactic Indicators and Sentential Contexts 4 Dec 2019 · 1 repository · arXiv:1912.01858
-
A Comparative Study of Pretrained Language Models on Thai Social Text Categorization 3 Dec 2019 · 0 repositories · arXiv:1912.01580
-
TU Wien @ TREC Deep Learning '19 -- Simple Contextualization for Re-ranking 3 Dec 2019 · 1 repository · arXiv:1912.01385
-
Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events 2 Dec 2019 · 0 repositories · arXiv:1912.02615
-
BERT for Large-scale Video Segment Classification with Test-time Augmentation 2 Dec 2019 · 0 repositories · arXiv:1912.01127
-
BLiMP: The Benchmark of Linguistic Minimal Pairs for English 2 Dec 2019 · 4 repositories · arXiv:1912.00582Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Leveraging Contextual Embeddings for Detecting Diachronic Semantic Shift 2 Dec 2019 · 0 repositories · arXiv:1912.01072
-
Long Distance Relationships without Time Travel: Boosting the Performance of a Sparse Predictive Autoencoder in Sequence Modeling 2 Dec 2019 · 1 repository · arXiv:1912.01116
-
Multi-Scale Self-Attention for Text Classification 2 Dec 2019 · 0 repositories · arXiv:1912.00544
-
Neural Academic Paper Generation 2 Dec 2019 · 1 repository · arXiv:1912.01982
-
Solving Arithmetic Word Problems Automatically Using Transformer and Unambiguous Representations 2 Dec 2019 · 1 repository · arXiv:1912.00871
-
Fast and Accurate Stochastic Gradient Estimation 1 Dec 2019 · 1 repository
-
Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks 1 Dec 2019 · 0 repositories
-
Perceiving the arrow of time in autoregressive motion 1 Dec 2019 · 0 repositories
-
Inducing Relational Knowledge from BERT 28 Nov 2019 · 0 repositories · arXiv:1911.12753
-
Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition 28 Nov 2019 · 0 repositories · arXiv:1911.12487
-
Automatic Generation of Headlines for Online Math Questions 27 Nov 2019 · 1 repository · arXiv:1912.00839
-
DeFINE: DEep Factorized INput Token Embeddings for Neural Sequence Modeling 27 Nov 2019 · 1 repository · arXiv:1911.12385