Methods › General › Attention Modules › Multi-Head Attention › Papers, page 185
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 185 of 249: papers 18,401 to 18,500 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Self-supervised clarification question generation for ambiguous multi-turn conversation 17 Dec 2021 · 0 repositories
-
Towards End-to-End Image Compression and Analysis with Transformers 17 Dec 2021 · 1 repository · arXiv:2112.09300
-
Towards Faithful Personalized Response Selection in Retrieval Based Dialog Systems 17 Dec 2021 · 0 repositories
-
UniMiSS: Universal Medical Self-Supervised Learning via Breaking Dimensionality Barrier 17 Dec 2021 · 1 repository · arXiv:2112.09356
-
WebGPT: Browser-assisted question-answering with human feedback 17 Dec 2021 · 2 repositories · arXiv:2112.09332
-
Your Answer is Incorrect... Would you like to know why? Introducing a Bilingual Short Answer Feedback Dataset 17 Dec 2021 · 0 repositories
-
An Empirical Study on Transfer Learning for Privilege Review 16 Dec 2021 · 0 repositories · arXiv:2112.08606
-
Block-Skim: Efficient Question Answering for Transformer 16 Dec 2021 · 1 repository · arXiv:2112.08560
-
Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge 16 Dec 2021 · 3 repositories · arXiv:2112.08619
-
Knowledge-Augmented Language Models for Cause-Effect Relation Classification 16 Dec 2021 · 1 repository · arXiv:2112.08615Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1,500+ Language Pairs 16 Dec 2021 · 1 repository · arXiv:2112.08804
-
Does Pre-training Induce Systematic Inference? How Masked Language Models Acquire Commonsense Knowledge 16 Dec 2021 · 0 repositories · arXiv:2112.08583
-
DProST: Dynamic Projective Spatial Transformer Network for 6D Pose Estimation 16 Dec 2021 · 1 repository · arXiv:2112.08775
-
DREAM: Improving Situational QA by First Elaborating the Situation 16 Dec 2021 · 1 repository · arXiv:2112.08656Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Efficient Hierarchical Domain Adaptation for Pretrained Language Models 16 Dec 2021 · 1 repository · arXiv:2112.08786
-
Few-Shot Semantic Parsing with Language Models Trained On Code 16 Dec 2021 · 0 repositories · arXiv:2112.08696
-
How to augment your ViTs? Consistency loss and StyleAug, a random style transfer augmentation 16 Dec 2021 · 0 repositories · arXiv:2112.09260
-
KAT: A Knowledge Augmented Transformer for Vision-and-Language 16 Dec 2021 · 1 repository · arXiv:2112.08614
-
Knowledge-enhanced Session-based Recommendation with Temporal Transformer 16 Dec 2021 · 0 repositories · arXiv:2112.08745
-
Learning Bounded Context-Free-Grammar via LSTM and the Transformer:Difference and Explanations 16 Dec 2021 · 1 repository · arXiv:2112.09174Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Learning Rich Representation of Keyphrases from Text 16 Dec 2021 · 1 repository · arXiv:2112.08547
-
Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds 16 Dec 2021 · 1 repository · arXiv:2112.08879
-
Multivariate Realized Volatility Forecasting with Graph Neural Network 16 Dec 2021 · 0 repositories · arXiv:2112.09015
-
Reconsidering the Past: Optimizing Hidden States in Language Models 16 Dec 2021 · 0 repositories · arXiv:2112.08653
-
Reframing Human-AI Collaboration for Generating Free-Text Explanations 16 Dec 2021 · 1 repository · arXiv:2112.08674Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense Reasoning 16 Dec 2021 · 0 repositories · arXiv:2112.08587
-
Trading with the Momentum Transformer: An Intelligent and Interpretable Architecture 16 Dec 2021 · 3 repositories · arXiv:2112.08534
-
TransZero++: Cross Attribute-Guided Transformer for Zero-Shot Learning 16 Dec 2021 · 1 repository · arXiv:2112.08643
-
Trees in transformers: a theoretical analysis of the Transformer's ability to represent trees 16 Dec 2021 · 0 repositories · arXiv:2112.11913
-
3D Question Answering 15 Dec 2021 · 0 repositories · arXiv:2112.08359
-
AllWOZ: Towards Multilingual Task-Oriented Dialog Systems for All 15 Dec 2021 · 0 repositories · arXiv:2112.08333
-
Applying SoftTriple Loss for Supervised Language Model Fine Tuning 15 Dec 2021 · 0 repositories · arXiv:2112.08462
-
Dense Video Captioning Using Unsupervised Semantic Information 15 Dec 2021 · 1 repository · arXiv:2112.08455
-
Fine-Tuning Large Neural Language Models for Biomedical Natural Language Processing 15 Dec 2021 · 0 repositories · arXiv:2112.07869
-
Is "My Favorite New Movie" My Favorite Movie? Probing the Understanding of Recursive Noun Phrases 15 Dec 2021 · 1 repository · arXiv:2112.08326
-
Learning to Transpile AMR into SPARQL 15 Dec 2021 · 0 repositories · arXiv:2112.07877
-
Lesan -- Machine Translation for Low Resource Languages 15 Dec 2021 · 0 repositories · arXiv:2112.08191
-
LongT5: Efficient Text-To-Text Transformer for Long Sequences 15 Dec 2021 · 4 repositories · arXiv:2112.07916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Named entity recognition architecture combining contextual and global features 15 Dec 2021 · 1 repository · arXiv:2112.08033
-
One size does not fit all: Investigating strategies for differentially-private learning across NLP tasks 15 Dec 2021 · 1 repository · arXiv:2112.08159Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
One System to Rule them All: a Universal Intent Recognition System for Customer Service Chatbots 15 Dec 2021 · 0 repositories · arXiv:2112.08261
-
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling 15 Dec 2021 · 1 repository · arXiv:2112.07874
-
SeqFormer: Sequential Transformer for Video Instance Segmentation 15 Dec 2021 · 2 repositories · arXiv:2112.08275
-
SPTS: Single-Point Text Spotting 15 Dec 2021 · 1 repository · arXiv:2112.07917
-
Tracing Text Provenance via Context-Aware Lexical Substitution 15 Dec 2021 · 0 repositories · arXiv:2112.07873
-
Vision Transformer Based Video Hashing Retrieval for Tracing the Source of Fake Videos 15 Dec 2021 · 0 repositories · arXiv:2112.08117
-
ACE-BERT: Adversarial Cross-modal Enhanced BERT for E-commerce Retrieval 14 Dec 2021 · 0 repositories · arXiv:2112.07209
-
AdaViT: Adaptive Tokens for Efficient Vision Transformer 14 Dec 2021 · 1 repository · arXiv:2112.07658Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Building on Huang et al. GlossBERT for Word Sense Disambiguation 14 Dec 2021 · 0 repositories · arXiv:2112.07089
-
Classifying Emails into Human vs Machine Category 14 Dec 2021 · 0 repositories · arXiv:2112.07742
-
CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising 14 Dec 2021 · 0 repositories · arXiv:2112.07515
-
Epigenomic language models powered by Cerebras 14 Dec 2021 · 0 repositories · arXiv:2112.07571
-
From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression 14 Dec 2021 · 2 repositories · arXiv:2112.07198
-
Geometry-Contrastive Transformer for Generalized 3D Pose Transfer 14 Dec 2021 · 1 repository · arXiv:2112.07374
-
Improving Compositional Generalization with Latent Structure and Data Augmentation 14 Dec 2021 · 2 repositories · arXiv:2112.07610
-
Improving Hybrid CTC/Attention End-to-end Speech Recognition with Pretrained Acoustic and Language Model 14 Dec 2021 · 0 repositories · arXiv:2112.07254
-
Measuring Fairness with Biased Rulers: A Survey on Quantifying Biases in Pretrained Language Models 14 Dec 2021 · 1 repository · arXiv:2112.07447
-
Temporal Transformer Networks with Self-Supervision for Action Recognition 14 Dec 2021 · 0 repositories · arXiv:2112.07338
-
Text Classification Models for Form Entity Linking 14 Dec 2021 · 1 repository · arXiv:2112.07443
-
Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text 14 Dec 2021 · 0 repositories · arXiv:2112.07074
-
5th Place Solution for VSPW 2021 Challenge 13 Dec 2021 · 0 repositories · arXiv:2112.06379
-
A Study on Token Pruning for ColBERT 13 Dec 2021 · 0 repositories · arXiv:2112.06540
-
Measuring Context-Word Biases in Lexical Semantic Datasets 13 Dec 2021 · 0 repositories · arXiv:2112.06733
-
Dependency Learning for Legal Judgment Prediction with a Unified Text-to-Text Transformer 13 Dec 2021 · 1 repository · arXiv:2112.06370
-
Do Data-based Curricula Work? 13 Dec 2021 · 0 repositories · arXiv:2112.06510
-
Embracing Single Stride 3D Object Detector with Sparse Transformer 13 Dec 2021 · 2 repositories · arXiv:2112.06375
-
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 13 Dec 2021 · 0 repositories · arXiv:2112.06905
-
Hformer: Hybrid CNN-Transformer for Fringe Order Prediction in Phase Unwrapping of Fringe Projection 13 Dec 2021 · 0 repositories · arXiv:2112.06759
-
Keyphrase Generation Beyond the Boundaries of Title and Abstract 13 Dec 2021 · 1 repository · arXiv:2112.06776
-
Pedestrian Trajectory Prediction via Spatial Interaction Transformer Network 13 Dec 2021 · 0 repositories · arXiv:2112.06624
-
Roof-Transformer: Divided and Joined Understanding with Knowledge Enhancement 13 Dec 2021 · 0 repositories · arXiv:2112.06736
-
Improving Sequential Recommendations via Bidirectional Temporal Data Augmentation with Pre-training 13 Dec 2021 · 1 repository · arXiv:2112.06460
-
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models 13 Dec 2021 · 1 repository · arXiv:2112.06598Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Implicit Transformer Network for Screen Content Image Continuous Super-Resolution 12 Dec 2021 · 1 repository · arXiv:2112.06174Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Improving Logical-Level Natural Language Generation with Topic-Conditioned Data Augmentation and Logical Form Generation 12 Dec 2021 · 0 repositories · arXiv:2112.06240
-
Improving Vision Transformers for Incremental Learning 12 Dec 2021 · 0 repositories · arXiv:2112.06103
-
Towards More Efficient Insertion Transformer with Fractional Positional Encoding 12 Dec 2021 · 1 repository · arXiv:2112.06295
-
COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality 11 Dec 2021 · 1 repository · arXiv:2112.05892
-
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition 10 Dec 2021 · 0 repositories · arXiv:2112.05820
-
Couplformer:Rethinking Vision Transformer with Coupling Attention Map 10 Dec 2021 · 1 repository · arXiv:2112.05425Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Deep ViT Features as Dense Visual Descriptors 10 Dec 2021 · 1 repository · arXiv:2112.05814Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Findings on Conversation Disentanglement 10 Dec 2021 · 0 repositories · arXiv:2112.05346
-
Multimodal Interactions Using Pretrained Unimodal Models for SIMMC 2.0 10 Dec 2021 · 1 repository · arXiv:2112.05328
-
Self-Supervised Transformers for fMRI representation 10 Dec 2021 · 2 repositories · arXiv:2112.05761
-
VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling 10 Dec 2021 · 0 repositories · arXiv:2112.05692
-
3D Medical Point Transformer: Introducing Convolution to Attention Networks for Medical Point Cloud Analysis 9 Dec 2021 · 1 repository · arXiv:2112.04863
-
A Bilingual, OpenWorld Video Text Dataset and End-to-end Video Text Spotter with Transformer 9 Dec 2021 · 3 repositories · arXiv:2112.04888Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 9 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
Detecting potentially harmful and protective suicide-related content on twitter: A machine learning approach 9 Dec 2021 · 2 repositories · arXiv:2112.04796
-
Extending AdamW by Leveraging Its Second Moment and Magnitude 9 Dec 2021 · 0 repositories · arXiv:2112.06125
-
Fast Point Transformer 9 Dec 2021 · 1 repository · arXiv:2112.04702Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
From Scattered Sources to Comprehensive Technology Landscape: A Recommendation-based Retrieval Approach 9 Dec 2021 · 0 repositories · arXiv:2112.04810
-
Injecting Semantic Concepts into End-to-End Image Captioning 9 Dec 2021 · 1 repository · arXiv:2112.05230Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
PE-former: Pose Estimation Transformer 9 Dec 2021 · 1 repository · arXiv:2112.04981
-
Recurrent Glimpse-based Decoder for Detection with Transformer 9 Dec 2021 · 1 repository · arXiv:2112.04632
-
Semantic Search as Extractive Paraphrase Span Detection 9 Dec 2021 · 1 repository · arXiv:2112.04886
-
Semi-Supervised Medical Image Segmentation via Cross Teaching between CNN and Transformer 9 Dec 2021 · 1 repository · arXiv:2112.04894Syntology official (archive's flag): 1 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Towards Neural Functional Program Evaluation 9 Dec 2021 · 0 repositories · arXiv:2112.04630
-
Garment4D: Garment Reconstruction from Point Cloud Sequences 8 Dec 2021 · 1 repository · arXiv:2112.04159Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Improving language models by retrieving from trillions of tokens 8 Dec 2021 · 2 repositories · arXiv:2112.04426Syntology 16 ran (of which 5 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 23 harvested samples) · 3 pointer-only (licence)
-
JABER and SABER: Junior and Senior Arabic BERt 8 Dec 2021 · 1 repository · arXiv:2112.04329