Methods › General › Attention Modules › Multi-Head Attention › Papers, page 220
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 220 of 249: papers 21,901 to 22,000 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
NIT COVID-19 at WNUT-2020 Task 2: Deep Learning Model RoBERTa for Identify Informative COVID-19 English Tweets 11 Nov 2020 · 0 repositories · arXiv:2011.05551
-
Recognizing More Emotions with Less Data Using Self-supervised Transfer Learning 11 Nov 2020 · 0 repositories · arXiv:2011.05585
-
Towards Semi-Supervised Semantics Understanding from Speech 11 Nov 2020 · 0 repositories · arXiv:2011.06195
-
TERMCast: Temporal Relation Modeling for Effective Urban Flow Forecasting 11 Nov 2020 · 0 repositories · arXiv:2011.05554
-
A Systematic Comparison of Encrypted Machine Learning Solutions for Image Classification 10 Nov 2020 · 0 repositories · arXiv:2011.05296
-
Detecting Social Media Manipulation in Low-Resource Languages 10 Nov 2020 · 0 repositories · arXiv:2011.05367
-
E.T.: Entity-Transformers. Coreference augmented Neural Language Model for richer mention representations via Entity-Transformer blocks 10 Nov 2020 · 0 repositories · arXiv:2011.05431
-
UmBERTo-MTSA @ AcCompl-It: Improving Complexity and Acceptability Prediction with Multi-task Learning on Self-Supervised Annotations 10 Nov 2020 · 1 repository · arXiv:2011.05197
-
When Do You Need Billions of Words of Pretraining Data? 10 Nov 2020 · 1 repository · arXiv:2011.04946
-
Bangla Text Classification using Transformers 9 Nov 2020 · 1 repository · arXiv:2011.04446
-
BERT-JAM: Boosting BERT-Enhanced Neural Machine Translation with Joint Attention 9 Nov 2020 · 0 repositories · arXiv:2011.04266
-
Positional Artefacts Propagate Through Masked Language Model Embeddings 9 Nov 2020 · 0 repositories · arXiv:2011.04393
-
CxGBERT: BERT meets Construction Grammar 9 Nov 2020 · 1 repository · arXiv:2011.04134
-
EstBERT: A Pretrained Language-Specific BERT for Estonian 9 Nov 2020 · 0 repositories · arXiv:2011.04784
-
Language Through a Prism: A Spectral Approach for Multiscale Language Representations 9 Nov 2020 · 1 repository · arXiv:2011.04823Syntology 0 ran · 5 unverified (of 5 harvested samples)
-
MAGNeto: An Efficient Deep Learning Method for the Extractive Tags Summarization Problem 9 Nov 2020 · 1 repository · arXiv:2011.04349
-
VisBERT: Hidden-State Visualizations for Transformers 9 Nov 2020 · 1 repository · arXiv:2011.04507
-
Adapting a Language Model for Controlled Affective Text Generation 8 Nov 2020 · 1 repository · arXiv:2011.04000Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages 8 Nov 2020 · 1 repository
-
Long Range Arena: A Benchmark for Efficient Transformers 8 Nov 2020 · 5 repositories · arXiv:2011.04006Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Stochastic Attention Head Removal: A simple and effective method for improving Transformer Based ASR Models 8 Nov 2020 · 5 repositories · arXiv:2011.04004Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Know What You Don't Need: Single-Shot Meta-Pruning for Attention Heads 7 Nov 2020 · 0 repositories · arXiv:2011.03770
-
Rethinking the Value of Transformer Components 7 Nov 2020 · 0 repositories · arXiv:2011.03803
-
SeqGenSQL -- A Robust Sequence Generation Model for Structured Query Language 7 Nov 2020 · 2 repositories · arXiv:2011.03836
-
From Dataset Recycling to Multi-Property Extraction and Beyond 6 Nov 2020 · 1 repository · arXiv:2011.03228
-
Highly Available Data Parallel ML training on Mesh Networks 6 Nov 2020 · 0 repositories · arXiv:2011.03605
-
Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis 6 Nov 2020 · 0 repositories · arXiv:2011.05161
-
Semi-Supervised Low-Resource Style Transfer of Indonesian Informal to Formal Language with Iterative Forward-Translation 6 Nov 2020 · 1 repository · arXiv:2011.03286
-
BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers 5 Nov 2020 · 0 repositories · arXiv:2011.02678
-
CODER: Knowledge infused cross-lingual medical term embedding for term normalization 5 Nov 2020 · 1 repository · arXiv:2011.02947
-
Competence-Level Prediction and Resume & Job Description Matching Using Context-Aware Transformer Models 5 Nov 2020 · 0 repositories · arXiv:2011.02998
-
NUAA-QMUL at SemEval-2020 Task 8: Utilizing BERT and DenseNet for Internet Meme Emotion Analysis 5 Nov 2020 · 1 repository · arXiv:2011.02788
-
Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages 4 Nov 2020 · 1 repository · arXiv:2011.02323
-
Investigating Novel Verb Learning in BERT: Selectional Preference Classes and Alternation-Based Syntactic Generalization 4 Nov 2020 · 1 repository · arXiv:2011.02417
-
MTLB-STRUCT @PARSEME 2020: Capturing Unseen Multiword Expressions Using Multi-task Learning and Pre-trained Masked Language Models 4 Nov 2020 · 1 repository · arXiv:2011.02541
-
Optimizing Transformer for Low-Resource Neural Machine Translation 4 Nov 2020 · 0 repositories · arXiv:2011.02266
-
Probing Multilingual BERT for Genetic and Typological Signals 4 Nov 2020 · 0 repositories · arXiv:2011.02070
-
Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech 4 Nov 2020 · 0 repositories · arXiv:2011.02252
-
BioNerFlair: biomedical named entity recognition using flair embedding and sequence tagger 3 Nov 2020 · 1 repository · arXiv:2011.01504
-
CharBERT: Character-aware Pre-trained Language Model 3 Nov 2020 · 1 repository · arXiv:2011.01513Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 2 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 3 pointer-only (licence)
-
Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image Wordings 3 Nov 2020 · 1 repository · arXiv:2011.01565
-
Finding Friends and Flipping Frenemies: Automatic Paraphrase Dataset Augmentation Using Graph Theory 3 Nov 2020 · 1 repository · arXiv:2011.01856Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Generating Synthetic Data for Task-Oriented Semantic Parsing with Hierarchical Representations 3 Nov 2020 · 0 repositories · arXiv:2011.02050
-
Improving RNN transducer with normalized jointer network 3 Nov 2020 · 0 repositories · arXiv:2011.01576
-
Sound Natural: Content Rephrasing in Dialog Systems 3 Nov 2020 · 1 repository · arXiv:2011.01993
-
Tabular Transformers for Modeling Multivariate Time Series 3 Nov 2020 · 1 repository · arXiv:2011.01843Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection 3 Nov 2020 · 1 repository · arXiv:2011.01612
-
A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English 2 Nov 2020 · 1 repository · arXiv:2011.00960
-
ABNIRML: Analyzing the Behavior of Neural IR Models 2 Nov 2020 · 2 repositories · arXiv:2011.00696Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Influence Patterns for Explaining Information Flow in BERT 2 Nov 2020 · 0 repositories · arXiv:2011.00740
-
Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation 2 Nov 2020 · 1 repository · arXiv:2011.00747
-
How Far Does BERT Look At:Distance-based Clustering and Analysis of BERT′s Attention 2 Nov 2020 · 0 repositories · arXiv:2011.00943
-
Introducing various Semantic Models for Amharic: Experimentation and Evaluation with multiple Tasks and Datasets 2 Nov 2020 · 1 repository · arXiv:2011.01154
-
On the Sentence Embeddings from Pre-trained Language Models 2 Nov 2020 · 3 repositories · arXiv:2011.05864Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Point Transformer 2 Nov 2020 · 2 repositories · arXiv:2011.00931Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
QMUL-SDS @ SardiStance: Leveraging Network Interactions to Boost Performance on Stance Detection using Knowledge Graphs 2 Nov 2020 · 0 repositories · arXiv:2011.01181
-
A Semantics-based Approach to Disclosure Classification in User-Generated Online Content 1 Nov 2020 · 0 repositories
-
A structure-enhanced graph convolutional network for sentiment analysis 1 Nov 2020 · 0 repositories
-
Active Learning Approaches to Enhancing Neural Machine Translation 1 Nov 2020 · 0 repositories
-
Approximation of Response Knowledge Retrieval in Knowledge-grounded Dialogue Generation 1 Nov 2020 · 1 repository
-
CHIME: Cross-passage Hierarchical Memory Network for Generative Review Question Answering 1 Nov 2020 · 1 repository · arXiv:2011.00519
-
ConceptBert: Concept-Aware Representation for Visual Question Answering 1 Nov 2020 · 2 repositories
-
Context Analysis for Pre-trained Masked Language Models 1 Nov 2020 · 0 repositories
-
COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning 1 Nov 2020 · 1 repository · arXiv:2011.00597Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Cross-Lingual Training of Neural Models for Document Ranking 1 Nov 2020 · 0 repositories
-
Decoding Language Spatial Relations to 2D Spatial Arrangements 1 Nov 2020 · 1 repository
-
Enhancing Generalization in Natural Language Inference by Syntax 1 Nov 2020 · 1 repository
-
exBERT: Extending Pre-trained Models with Domain-specific Vocabulary Under Constrained Training Resources 1 Nov 2020 · 0 repositories
-
Factorized Transformer for Multi-Domain Neural Machine Translation 1 Nov 2020 · 0 repositories
-
Hate-Speech and Offensive Language Detection in Roman Urdu 1 Nov 2020 · 0 repositories
-
HUJI-KU at MRP 2020: Two Transition-based Neural Parsers 1 Nov 2020 · 0 repositories
-
Integrating Task Specific Information into Pretrained Language Models for Low Resource Fine Tuning 1 Nov 2020 · 1 repository
-
Investigation of BERT Model on Biomedical Relation Extraction Based on Revised Fine-tuning Mechanism 1 Nov 2020 · 0 repositories · arXiv:2011.00398
-
Learning to ground medical text in a 3D human atlas 1 Nov 2020 · 1 repository
-
LIMIT-BERT : Linguistics Informed Multi-Task BERT 1 Nov 2020 · 1 repository
-
Making Information Seeking Easier: An Improved Pipeline for Conversational Search 1 Nov 2020 · 0 repositories
-
Modeling Intra and Inter-modality Incongruity for Multi-Modal Sarcasm Detection 1 Nov 2020 · 0 repositories
-
Multi\^2OIE: Multilingual Open Information Extraction Based on Multi-Head Attention with BERT 1 Nov 2020 · 1 repository
-
Optimizing Word Segmentation for Downstream Task 1 Nov 2020 · 1 repository
-
Predicting Responses to Psychological Questionnaires from Participants' Social Media Posts and Question Text Embeddings 1 Nov 2020 · 0 repositories
-
Representation Learning for Type-Driven Composition 1 Nov 2020 · 1 repository
-
Seeing Both the Forest and the Trees: Multi-head Attention for Joint Classification on Different Compositional Levels 1 Nov 2020 · 1 repository · arXiv:2011.00470
-
SMRT Chatbots: Improving Non-Task-Oriented Dialog with Simulated Multiple Reference Training 1 Nov 2020 · 0 repositories · arXiv:2011.00547
-
Social Chemistry 101: Learning to Reason about Social and Moral Norms 1 Nov 2020 · 2 repositories · arXiv:2011.00620
-
The Amazing World of Neural Language Generation 1 Nov 2020 · 0 repositories
-
The RELX Dataset and Matching the Multilingual Blanks for Cross-Lingual Relation Classification 1 Nov 2020 · 1 repository
-
Towards Zero-Shot Conditional Summarization with Adaptive Multi-Task Fine-Tuning 1 Nov 2020 · 1 repository
-
Transformer-based Multi-Aspect Modeling for Multi-Aspect Multi-Sentiment Analysis 1 Nov 2020 · 0 repositories · arXiv:2011.00476
-
Visually-Grounded Planning without Vision: Language Models Infer Detailed Plans from High-level Instructions 1 Nov 2020 · 1 repository
-
Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution 31 Oct 2020 · 1 repository · arXiv:2011.00245
-
Neural Coreference Resolution for Arabic 31 Oct 2020 · 1 repository · arXiv:2011.00286
-
Understanding Pre-trained BERT for Aspect-based Sentiment Analysis 31 Oct 2020 · 2 repositories · arXiv:2011.00169
-
A Sui Generis QA Approach using RoBERTa for Adverse Drug Event Identification 30 Oct 2020 · 0 repositories · arXiv:2011.00057
-
Generating Radiology Reports via Memory-driven Transformer 30 Oct 2020 · 2 repositories · arXiv:2010.16056Syntology official (archive's flag): 1 ran · 26 ran (of which 13 constructed an object rather than computing a result; 20 with no instrument failure: 3 honoured, 3 violated, 14 with no contract checked; 6 where Syntology's instrument failed) · 5 unverified (of 31 harvested samples) · 25 pointer-only (licence)
-
Improving Dialogue Breakdown Detection with Semi-Supervised Learning 30 Oct 2020 · 0 repositories · arXiv:2011.00136
-
Semantic Labeling Using a Deep Contextualized Language Model 30 Oct 2020 · 1 repository · arXiv:2010.16037
-
SLM: Learning a Discourse Language Representation with Sentence Unshuffling 30 Oct 2020 · 0 repositories · arXiv:2010.16249
-
Target Word Masking for Location Metonymy Resolution 30 Oct 2020 · 1 repository · arXiv:2010.16097
-
Topic-Preserving Synthetic News Generation: An Adversarial Deep Reinforcement Learning Approach 30 Oct 2020 · 0 repositories · arXiv:2010.16324
-
VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation 30 Oct 2020 · 1 repository · arXiv:2010.16046