Methods › General › Attention Modules › Multi-Head Attention › Papers, page 197
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 197 of 249: papers 19,601 to 19,700 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LegaLMFiT: Efficient Short Legal Text Classification with LSTM Language Model Pre-Training 2 Sep 2021 · 0 repositories · arXiv:2109.00993
-
So Cloze yet so Far: N400 Amplitude is Better Predicted by Distributional Information than Human Predictability Judgements 2 Sep 2021 · 0 repositories · arXiv:2109.01226
-
Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation 2 Sep 2021 · 0 repositories · arXiv:2109.01048
-
Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction 1 Sep 2021 · 1 repository · arXiv:2109.00512
-
CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations 1 Sep 2021 · 1 repository · arXiv:2109.00181Syntology official (archive's flag): 13 ran · 15 ran (of which 2 constructed an object rather than computing a result; 15 with no instrument failure: 1 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples) · 6 pointer-only (licence)
-
DILBERT: Customized Pre-Training for Domain Adaptation withCategory Shift, with an Application to Aspect Extraction 1 Sep 2021 · 1 repository · arXiv:2109.00571
-
Exploring deep learning methods for recognizing rare diseases and their clinical manifestations from texts 1 Sep 2021 · 2 repositories · arXiv:2109.00343
-
Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech 1 Sep 2021 · 0 repositories · arXiv:2109.00591
-
∞-former: Infinite Memory Transformer 1 Sep 2021 · 1 repository · arXiv:2109.00301
-
OptAGAN: Entropy-based finetuning on text VAE-GAN 1 Sep 2021 · 1 repository · arXiv:2109.00239
-
Searching for Efficient Multi-Stage Vision Transformers 1 Sep 2021 · 1 repository · arXiv:2109.00642
-
Stochastic Transformer Networks with Linear Competing Units: Application to end-to-end SL Translation 1 Sep 2021 · 1 repository · arXiv:2109.13318Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Towards Improving Adversarial Training of NLP Models 1 Sep 2021 · 1 repository · arXiv:2109.00544Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
AraT5: Text-to-Text Transformers for Arabic Language Generation 31 Aug 2021 · 1 repository · arXiv:2109.12068
-
Automated Mining of Leaderboards for Empirical AI Research 31 Aug 2021 · 1 repository · arXiv:2109.13089
-
Effectiveness of Deep Networks in NLP using BiDAF as an example architecture 31 Aug 2021 · 0 repositories · arXiv:2109.00074
-
Enjoy the Salience: Towards Better Transformer-based Faithful Explanations with Word Salience 31 Aug 2021 · 1 repository · arXiv:2108.13759
-
How Does Adversarial Fine-Tuning Benefit BERT? 31 Aug 2021 · 0 repositories · arXiv:2108.13602
-
SANSformers: Self-Supervised Forecasting in Electronic Health Records with Attention-Free Models 31 Aug 2021 · 0 repositories · arXiv:2108.13672
-
MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics 31 Aug 2021 · 4 repositories · arXiv:2109.00110Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Monolingual versus Multilingual BERTology for Vietnamese Extractive Multi-Document Summarization 31 Aug 2021 · 0 repositories · arXiv:2108.13741
-
Sense representations for Portuguese: experiments with sense embeddings and deep neural language models 31 Aug 2021 · 0 repositories · arXiv:2109.00025
-
Task-Oriented Dialogue System as Natural Language Generation 31 Aug 2021 · 1 repository · arXiv:2108.13679Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP 30 Aug 2021 · 1 repository · arXiv:2108.13002
-
Exploring and Improving Mobile Level Vision Transformers 30 Aug 2021 · 0 repositories · arXiv:2108.13015
-
From General to Specific: Informative Scene Graph Generation via Balance Adjustment 30 Aug 2021 · 1 repository · arXiv:2108.13129Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Improving Query Representations for Dense Retrieval with Pseudo Relevance Feedback 30 Aug 2021 · 2 repositories · arXiv:2108.13454Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Knowledge Base Completion Meets Transfer Learning 30 Aug 2021 · 1 repository · arXiv:2108.13073Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Multi-Channel Transformer Transducer for Speech Recognition 30 Aug 2021 · 0 repositories · arXiv:2108.12953
-
On the Multilingual Capabilities of Very Large-Scale English Language Models 30 Aug 2021 · 2 repositories · arXiv:2108.13349
-
Scheduled Sampling Based on Decoding Steps for Neural Machine Translation 30 Aug 2021 · 1 repository · arXiv:2108.12963
-
Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning 30 Aug 2021 · 0 repositories · arXiv:2108.13032
-
Want To Reduce Labeling Cost? GPT-3 Can Help 30 Aug 2021 · 1 repository · arXiv:2108.13487
-
Analyzing and Mitigating Interference in Neural Architecture Search 29 Aug 2021 · 0 repositories · arXiv:2108.12821
-
NoiER: An Approach for Training more Reliable Fine-TunedDownstream Task Models 29 Aug 2021 · 0 repositories · arXiv:2110.02054
-
TCCT: Tightly-Coupled Convolutional Transformer on Time Series Forecasting 29 Aug 2021 · 2 repositories · arXiv:2108.12784
-
AMMASurv: Asymmetrical Multi-Modal Attention for Accurate Survival Analysis with Whole Slide Images and Gene Expression Data 28 Aug 2021 · 0 repositories · arXiv:2108.12565
-
DKM: Differentiable K-Means Clustering Layer for Neural Network Compression 28 Aug 2021 · 0 repositories · arXiv:2108.12659
-
GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer 28 Aug 2021 · 1 repository · arXiv:2108.12630Syntology official (archive's flag): 20 ran · 22 ran (of which 2 constructed an object rather than computing a result; 8 with no instrument failure: 3 honoured, 0 violated, 5 with no contract checked; 14 where Syntology's instrument failed) · 5 unverified (of 27 harvested samples) · 2 pointer-only (licence)
-
HeadlineCause: A Dataset of News Headlines for Detecting Causalities 28 Aug 2021 · 1 repository · arXiv:2108.12626
-
Towards Fine-grained Image Classification with Generative Adversarial Networks and Facial Landmark Detection 28 Aug 2021 · 1 repository · arXiv:2109.00891
-
Automatic Text Evaluation through the Lens of Wasserstein Barycenters 27 Aug 2021 · 2 repositories · arXiv:2108.12463
-
Dealing with Typos for BERT-based Passage Retrieval and Ranking 27 Aug 2021 · 2 repositories · arXiv:2108.12139
-
Evaluating the Robustness of Neural Language Models to Input Perturbations 27 Aug 2021 · 1 repository · arXiv:2108.12237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Lyra: A Benchmark for Turducken-Style Code Generation 27 Aug 2021 · 1 repository · arXiv:2108.12144
-
Query-Focused Extractive Summarisation for Finding Ideal Answers to Biomedical and COVID-19 Questions 27 Aug 2021 · 1 repository · arXiv:2108.12189
-
A Computational Approach to Measure Empathy and Theory-of-Mind from Written Texts 26 Aug 2021 · 1 repository · arXiv:2108.11810
-
A New Sentence Ordering Method Using BERT Pretrained Model 26 Aug 2021 · 0 repositories · arXiv:2108.11994
-
Can the Transformer Be Used as a Drop-in Replacement for RNNs in Text-Generating GANs? 26 Aug 2021 · 0 repositories · arXiv:2108.12275
-
EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa 26 Aug 2021 · 1 repository · arXiv:2108.12009
-
Evaluating Transformer-based Semantic Segmentation Networks for Pathological Image Segmentation 26 Aug 2021 · 0 repositories · arXiv:2108.11993
-
DeepGene Transformer: Transformer for the gene expression-based classification of cancer subtypes 26 Aug 2021 · 0 repositories · arXiv:2108.11833
-
LayoutReader: Pre-training of Text and Layout for Reading Order Detection 26 Aug 2021 · 1 repository · arXiv:2108.11591
-
Reiterative Domain Aware Multi-Target Adaptation 26 Aug 2021 · 0 repositories · arXiv:2109.00919
-
Rethinking Why Intermediate-Task Fine-Tuning Works 26 Aug 2021 · 1 repository · arXiv:2108.11696
-
Shifted Chunk Transformer for Spatio-Temporal Representational Learning 26 Aug 2021 · 0 repositories · arXiv:2108.11575
-
SLIM: Explicit Slot-Intent Mapping with BERT for Joint Multi-Intent Detection and Slot Filling 26 Aug 2021 · 1 repository · arXiv:2108.11711
-
The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers 26 Aug 2021 · 2 repositories · arXiv:2108.12284Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios 26 Aug 2021 · 3 repositories · arXiv:2108.11539
-
Multilingual Multi-Aspect Explainability Analyses on Machine Reading Comprehension Models 26 Aug 2021 · 1 repository · arXiv:2108.11574
-
CancerBERT: a BERT model for Extracting Breast Cancer Phenotypes from Electronic Health Records 25 Aug 2021 · 0 repositories · arXiv:2108.11303
-
Transformer for Single Image Super-Resolution 25 Aug 2021 · 1 repository · arXiv:2108.11084Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens 25 Aug 2021 · 1 repository · arXiv:2108.11193Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
On Approximate Nearest Neighbour Selection for Multi-Stage Dense Retrieval 25 Aug 2021 · 1 repository · arXiv:2108.11480
-
Ontology-Enhanced Slot Filling 25 Aug 2021 · 0 repositories · arXiv:2108.11275
-
What do pre-trained code models know about code? 25 Aug 2021 · 1 repository · arXiv:2108.11308
-
Auto-Parsing Network for Image Captioning and Visual Question Answering 24 Aug 2021 · 0 repositories · arXiv:2108.10568
-
Greenformers: Improving Computation and Memory Efficiency in Transformer Models via Low-Rank Approximation 24 Aug 2021 · 0 repositories · arXiv:2108.10808
-
sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification 24 Aug 2021 · 1 repository · arXiv:2108.10566
-
Towards Offensive Language Identification for Tamil Code-Mixed YouTube Comments and Posts 24 Aug 2021 · 1 repository · arXiv:2108.10939
-
Using BERT Encoding and Sentence-Level Language Model for Sentence Ordering 24 Aug 2021 · 0 repositories · arXiv:2108.10986
-
Weakly Supervised Cross-platform Teenager Detection with Adversarial BERT 24 Aug 2021 · 0 repositories · arXiv:2108.10619
-
CGEMs: A Metric Model for Automatic Code Generation using GPT-3 23 Aug 2021 · 0 repositories · arXiv:2108.10168
-
Deploying a BERT-based Query-Title Relevance Classifier in a Production System: a View from the Trenches 23 Aug 2021 · 0 repositories · arXiv:2108.10197
-
Improving 3D Object Detection with Channel-wise Transformer 23 Aug 2021 · 1 repository · arXiv:2108.10723Syntology official (archive's flag): 6 ran · 6 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples)
-
One TTS Alignment To Rule Them All 23 Aug 2021 · 3 repositories · arXiv:2108.10447
-
Query Embedding Pruning for Dense Retrieval 23 Aug 2021 · 1 repository · arXiv:2108.10341
-
Recurrent multiple shared layers in Depth for Neural Machine Translation 23 Aug 2021 · 0 repositories · arXiv:2108.10417
-
Regularizing Transformers With Deep Probabilistic Layers 23 Aug 2021 · 0 repositories · arXiv:2108.10764
-
Sarcasm Detection in Twitter -- Performance Impact while using Data Augmentation: Word Embeddings 23 Aug 2021 · 1 repository · arXiv:2108.09924
-
SwinIR: Image Restoration Using Swin Transformer 23 Aug 2021 · 9 repositories · arXiv:2108.10257Syntology official (archive's flag): 3 ran · 30 ran (of which 14 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 14 where Syntology's instrument failed) · 15 unverified (of 45 harvested samples) · 5 pointer-only (licence)
-
ZS-SLR: Zero-Shot Sign Language Recognition from RGB-D Videos 23 Aug 2021 · 0 repositories · arXiv:2108.10059
-
Guiding Query Position and Performing Similar Attention for Transformer-Based Detection Heads 22 Aug 2021 · 0 repositories · arXiv:2108.09691
-
Spatial Transformer Networks for Curriculum Learning 22 Aug 2021 · 0 repositories · arXiv:2108.09696
-
StarVQA: Space-Time Attention for Video Quality Assessment 22 Aug 2021 · 0 repositories · arXiv:2108.09635
-
Using Large Pre-Trained Models with Cross-Modal Attention for Multi-Modal Emotion Recognition 22 Aug 2021 · 0 repositories · arXiv:2108.09669
-
UzBERT: pretraining a BERT model for Uzbek 22 Aug 2021 · 0 repositories · arXiv:2108.09814
-
Construction material classification on imbalanced datasets using Vision Transformer (ViT) architecture 21 Aug 2021 · 0 repositories · arXiv:2108.09527
-
Approximate Bayesian Neural Doppler Imaging 20 Aug 2021 · 1 repository · arXiv:2108.09266
-
Convolutional Neural Network (CNN) vs Vision Transformer (ViT) for Digital Holography 20 Aug 2021 · 0 repositories · arXiv:2108.09147
-
Extracting Radiological Findings With Normalized Anatomical Information Using a Span-Based BERT Relation Extraction Model 20 Aug 2021 · 0 repositories · arXiv:2108.09211
-
Fastformer: Additive Attention Can Be All You Need 20 Aug 2021 · 13 repositories · arXiv:2108.09084Syntology official: no sample here; runs from other or unrecorded repositories · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Frozen Pretrained Transformers for Neural Sign Language Translation 20 Aug 2021 · 1 repository
-
MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition 20 Aug 2021 · 0 repositories · arXiv:2108.09322
-
One Chatbot Per Person: Creating Personalized Chatbots based on Implicit User Profiles 20 Aug 2021 · 1 repository · arXiv:2108.09355
-
Pre-training for Ad-hoc Retrieval: Hyperlink is Also You Need 20 Aug 2021 · 1 repository · arXiv:2108.09346
-
Semantic Communication with Adaptive Universal Transformer 20 Aug 2021 · 0 repositories · arXiv:2108.09119
-
Smart Bird: Learnable Sparse Attention for Efficient and Effective Transformer 20 Aug 2021 · 0 repositories · arXiv:2108.09193
-
Trans4Trans: Efficient Transformer for Transparent Object and Semantic Scene Segmentation in Real-World Navigation Assistance 20 Aug 2021 · 1 repository · arXiv:2108.09174
-
Uncertainties and output feedback in rollout event-triggered control 20 Aug 2021 · 0 repositories · arXiv:2108.09125