Methods › General › Attention Modules › Multi-Head Attention › Papers, page 212
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 212 of 249: papers 21,101 to 21,200 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SpecTr: Spectral Transformer for Hyperspectral Pathology Image Segmentation 5 Mar 2021 · 1 repository · arXiv:2103.03604
-
CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation 4 Mar 2021 · 1 repository · arXiv:2103.03024
-
End-to-end acoustic modelling for phone recognition of young readers 4 Mar 2021 · 0 repositories · arXiv:2103.02899
-
Hardware Acceleration of Fully Quantized BERT for Efficient Natural Language Processing 4 Mar 2021 · 0 repositories · arXiv:2103.02800
-
The Transformer Network for the Traveling Salesman Problem 4 Mar 2021 · 1 repository · arXiv:2103.03012
-
Few-shot Learning for Slot Tagging with Attentive Relational Network 3 Mar 2021 · 0 repositories · arXiv:2103.02333
-
Natural Language Understanding for Argumentative Dialogue Systems in the Opinion Building Domain 3 Mar 2021 · 0 repositories · arXiv:2103.02691
-
University of Copenhagen Participation in TREC Health Misinformation Track 2020 3 Mar 2021 · 0 repositories · arXiv:2103.02462
-
Video Sentiment Analysis with Bimodal Information-augmented Multi-Head Attention 3 Mar 2021 · 0 repositories · arXiv:2103.02362
-
A Minimalist Dataset for Systematic Generalization of Perception, Syntax, and Semantics 2 Mar 2021 · 0 repositories · arXiv:2103.01403
-
Disentangling Syntax and Semantics in the Brain with Deep Networks 2 Mar 2021 · 0 repositories · arXiv:2103.01620
-
Dual Reinforcement-Based Specification Generation for Image De-Rendering 2 Mar 2021 · 0 repositories · arXiv:2103.01867
-
Hate Towards the Political Opponent: A Twitter Corpus Study of the 2020 US Elections on the Basis of Offensive Speech and Stance Detection 2 Mar 2021 · 0 repositories · arXiv:2103.01664
-
Probing Product Description Generation via Posterior Distillation 2 Mar 2021 · 0 repositories · arXiv:2103.01594
-
BERT-based knowledge extraction method of unstructured domain text 1 Mar 2021 · 0 repositories · arXiv:2103.00728
-
BERT based patent novelty search by training claims to their own description 1 Mar 2021 · 0 repositories · arXiv:2103.01126
-
Combat COVID-19 Infodemic Using Explainable Natural Language Processing Models 1 Mar 2021 · 0 repositories · arXiv:2103.00747
-
CrossMap Transformer: A Crossmodal Masked Path Transformer Using Double Back-Translation for Vision-and-Language Navigation 1 Mar 2021 · 0 repositories · arXiv:2103.00852
-
Generative Adversarial Transformers 1 Mar 2021 · 2 repositories · arXiv:2103.01209Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Long Document Summarization in a Low Resource Setting using Pretrained Language Models 1 Mar 2021 · 0 repositories · arXiv:2103.00751
-
NLP-CUET@DravidianLangTech-EACL2021: Investigating Visual and Textual Features to Identify Trolls from Multimodal Social Media Memes 28 Feb 2021 · 0 repositories · arXiv:2103.00466
-
NLP-CUET@DravidianLangTech-EACL2021: Offensive Language Detection from Multilingual Code-Mixed Text using Transformers 28 Feb 2021 · 1 repository · arXiv:2103.00455
-
NLP-CUET@LT-EDI-EACL2021: Multilingual Code-Mixed Hope Speech Detection using Cross-lingual Representation Learner 28 Feb 2021 · 1 repository · arXiv:2103.00464
-
COVID-19 Tweets Analysis through Transformer Language Models 27 Feb 2021 · 1 repository · arXiv:2103.00199
-
Generative Chemical Transformer: Neural Machine Learning of Molecular Geometric Structures from Chemical Language via Attention 27 Feb 2021 · 2 repositories · arXiv:2103.00213
-
Transformer in Transformer 27 Feb 2021 · 12 repositories · arXiv:2103.00112Syntology official (archive's flag): 4 ran · 16 ran (of which 12 constructed an object rather than computing a result; 15 with no instrument failure: 1 honoured, 1 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 24 harvested samples) · 5 pointer-only (licence)
-
Transformers with Competitive Ensembles of Independent Mechanisms 27 Feb 2021 · 0 repositories · arXiv:2103.00336
-
Multi-task transfer learning for finding actionable information from crisis-related messages on social media 26 Feb 2021 · 0 repositories · arXiv:2102.13395
-
BERT-based Acronym Disambiguation with Multiple Training Strategies 25 Feb 2021 · 0 repositories · arXiv:2103.00488
-
Emotion-Aware, Emotion-Agnostic, or Automatic: Corpus Creation Strategies to Obtain Cognitive Event Appraisal Annotations 25 Feb 2021 · 0 repositories · arXiv:2102.12858
-
LazyFormer: Self Attention with Lazy Update 25 Feb 2021 · 0 repositories · arXiv:2102.12702
-
LET: Linguistic Knowledge Enhanced Graph Transformer for Chinese Short Text Matching 25 Feb 2021 · 1 repository · arXiv:2102.12671
-
MixSpeech: Data Augmentation for Low-resource Automatic Speech Recognition 25 Feb 2021 · 0 repositories · arXiv:2102.12664
-
PharmKE: Knowledge Extraction Platform for Pharmaceutical Texts using Transfer Learning 25 Feb 2021 · 0 repositories · arXiv:2102.13139
-
Sentiment Analysis of Persian-English Code-mixed Texts 25 Feb 2021 · 1 repository · arXiv:2102.12700
-
From Universal Language Model to Downstream Task: Improving RoBERTa-Based Vietnamese Hate Speech Detection 24 Feb 2021 · 0 repositories · arXiv:2102.12162
-
Hopeful_Men@LT-EDI-EACL2021: Hope Speech Detection Using Indic Transliteration and Transformers 24 Feb 2021 · 0 repositories · arXiv:2102.12082
-
LRG at SemEval-2021 Task 4: Improving Reading Comprehension with Abstract Words using Augmentation, Linguistic Features and Voting 24 Feb 2021 · 1 repository · arXiv:2102.12255
-
NLRG at SemEval-2021 Task 5: Toxic Spans Detection Leveraging BERT-based Token Classification and Span Prediction Techniques 24 Feb 2021 · 1 repository · arXiv:2102.12254
-
PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen Domains 24 Feb 2021 · 1 repository · arXiv:2102.12206
-
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions 24 Feb 2021 · 11 repositories · arXiv:2102.12122Syntology official (archive's flag): 4 ran · 22 ran (of which 16 constructed an object rather than computing a result; 18 with no instrument failure: 2 honoured, 0 violated, 16 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 30 harvested samples) · 1 pointer-only (licence)
-
Task-Specific Pre-Training and Cross Lingual Transfer for Code-Switched Data 24 Feb 2021 · 0 repositories · arXiv:2102.12407
-
When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute 24 Feb 2021 · 1 repository · arXiv:2102.12459
-
Accurate Learning of Graph Representations with Graph Multiset Pooling 23 Feb 2021 · 1 repository · arXiv:2102.11533Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Deep Deformation Detail Synthesis for Thin Shell Models 23 Feb 2021 · 0 repositories · arXiv:2102.11541
-
Do Transformer Modifications Transfer Across Implementations and Applications? 23 Feb 2021 · 1 repository · arXiv:2102.11972
-
Minimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich Networks 23 Feb 2021 · 0 repositories · arXiv:2102.11479
-
Robust and Transferable Anomaly Detection in Log Data using Pre-Trained Language Models 23 Feb 2021 · 0 repositories · arXiv:2102.11570
-
VisualCheXbert: Addressing the Discrepancy Between Radiology Report Labels and Image Labels 23 Feb 2021 · 1 repository · arXiv:2102.11467Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Deepfake Video Detection Using Convolutional Vision Transformer 22 Feb 2021 · 1 repository · arXiv:2102.11126
-
Determination of Fault Location in Transmission Lines with Image Processing and Artificial Neural Networks 22 Feb 2021 · 0 repositories · arXiv:2102.11073
-
Conditional Positional Encodings for Vision Transformers 22 Feb 2021 · 2 repositories · arXiv:2102.10882
-
Evaluating Contextualized Language Models for Hungarian 22 Feb 2021 · 1 repository · arXiv:2102.10848
-
Few Shot Learning for Information Verification 22 Feb 2021 · 0 repositories · arXiv:2102.10956
-
Generating Human Readable Transcript for Automatic Speech Recognition with Pre-trained Language Model 22 Feb 2021 · 0 repositories · arXiv:2102.11114
-
MixUp Training Leads to Reduced Overfitting and Improved Calibration for the Transformer Architecture 22 Feb 2021 · 0 repositories · arXiv:2102.11402
-
Parallelizing Legendre Memory Unit Training 22 Feb 2021 · 2 repositories · arXiv:2102.11417
-
Position Information in Transformers: An Overview 22 Feb 2021 · 0 repositories · arXiv:2102.11090
-
RUBERT: A Bilingual Roman Urdu BERT Using Cross Lingual Transfer Learning 22 Feb 2021 · 0 repositories · arXiv:2102.11278
-
UniT: Multimodal Multitask Learning with a Unified Transformer 22 Feb 2021 · 1 repository · arXiv:2102.10772
-
Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks 22 Feb 2021 · 1 repository · arXiv:2102.10934
-
Medical Transformer: Gated Axial-Attention for Medical Image Segmentation 21 Feb 2021 · 2 repositories · arXiv:2102.10662
-
Pre-Training BERT on Arabic Tweets: Practical Considerations 21 Feb 2021 · 0 repositories · arXiv:2102.10684
-
Web-based Application for Detecting Indonesian Clickbait Headlines using IndoBERT 21 Feb 2021 · 0 repositories · arXiv:2102.10601
-
Multilingual Answer Sentence Reranking via Automatically Translated Data 20 Feb 2021 · 0 repositories · arXiv:2102.10250
-
Towards Accurate and Compact Architectures via Neural Architecture Transformer 20 Feb 2021 · 2 repositories · arXiv:2102.10301
-
Calibrate Before Use: Improving Few-Shot Performance of Language Models 19 Feb 2021 · 5 repositories · arXiv:2102.09690Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
Dialect Identification in Nuanced Arabic Tweets Using Farasa Segmentation and AraBERT 19 Feb 2021 · 0 repositories · arXiv:2102.09749
-
Latent Variable Sequential Set Transformers For Joint Multi-Agent Motion Prediction 19 Feb 2021 · 2 repositories · arXiv:2104.00563Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 17 harvested samples) · 2 pointer-only (licence)
-
Learning Dynamic BERT via Trainable Gate Variables and a Bi-modal Regularizer 19 Feb 2021 · 0 repositories · arXiv:2102.09727
-
Towards Emotion Recognition in Hindi-English Code-Mixed Data: A Transformer Based Approach 19 Feb 2021 · 1 repository · arXiv:2102.09943
-
Using Transformer based Ensemble Learning to classify Scientific Articles 19 Feb 2021 · 2 repositories · arXiv:2102.09991
-
Analysis Of Contextual and Non-Contextual Word Embedding Models For Hindi NER With Web Application For Data Collection 18 Feb 2021 · 1 repository
-
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer 18 Feb 2021 · 1 repository · arXiv:2102.09550
-
Quiz-Style Question Generation for News Stories 18 Feb 2021 · 2 repositories · arXiv:2102.09094
-
Training Large-Scale News Recommenders with Pretrained Language Models in the Loop 18 Feb 2021 · 1 repository · arXiv:2102.09268
-
UnibucKernel: Geolocating Swiss German Jodels Using Ensemble Learning 18 Feb 2021 · 0 repositories · arXiv:2102.09379
-
Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with 1/n Parameters 17 Feb 2021 · 3 repositories · arXiv:2102.08597Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Leveraging Query Resolution and Reading Comprehension for Conversational Passage Retrieval 17 Feb 2021 · 0 repositories · arXiv:2102.08795
-
SciDr at SDU-2020: IDEAS -- Identifying and Disambiguating Everyday Acronyms for Scientific Domain 17 Feb 2021 · 2 repositories · arXiv:2102.08818
-
TCN: Table Convolutional Network for Web Table Interpretation 17 Feb 2021 · 1 repository · arXiv:2102.09460
-
THEaiTRE 1.0: Interactive generation of theatre play scripts 17 Feb 2021 · 0 repositories · arXiv:2102.08892
-
COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining 16 Feb 2021 · 2 repositories · arXiv:2102.08473Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet 16 Feb 2021 · 1 repository · arXiv:2102.08036
-
A2-FPN for Semantic Segmentation of Fine-Resolution Remotely Sensed Images 16 Feb 2021 · 2 repositories · arXiv:2102.07997
-
GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training 16 Feb 2021 · 2 repositories · arXiv:2102.08098Syntology official (archive's flag): 6 ran · 7 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples) · 12 pointer-only (licence)
-
Have Attention Heads in BERT Learned Constituency Grammar? 16 Feb 2021 · 0 repositories · arXiv:2102.07926
-
Non-Autoregressive Text Generation with Pre-trained Language Models 16 Feb 2021 · 1 repository · arXiv:2102.08220
-
Revisiting Language Encoding in Learning Multilingual Representations 16 Feb 2021 · 1 repository · arXiv:2102.08357
-
TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models 16 Feb 2021 · 1 repository · arXiv:2102.07988Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
DOBF: A Deobfuscation Pre-Training Objective for Programming Languages 15 Feb 2021 · 2 repositories · arXiv:2102.07492
-
Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT 15 Feb 2021 · 0 repositories · arXiv:2102.07594
-
Improved Customer Transaction Classification using Semi-Supervised Knowledge Distillation 15 Feb 2021 · 0 repositories · arXiv:2102.07635
-
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm 15 Feb 2021 · 0 repositories · arXiv:2102.07350
-
The corruptive force of AI-generated advice 15 Feb 2021 · 0 repositories · arXiv:2102.07536
-
Translational Equivariance in Kernelizable Attention 15 Feb 2021 · 1 repository · arXiv:2102.07680
-
Within-Document Event Coreference with BERT-Based Contextualized Representations 15 Feb 2021 · 0 repositories · arXiv:2102.09600
-
indicnlp@kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository · arXiv:2102.07150
-
indicnlp@ kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository
-
Query-by-Example Keyword Spotting system using Multi-head Attention and Softtriple Loss 14 Feb 2021 · 0 repositories · arXiv:2102.07061