Methods › General › Attention Modules › Multi-Head Attention › Papers, page 183
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 183 of 249: papers 18,201 to 18,300 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
TEMPLATE: TempRel Classification Model Trained with Embedded Temporal Relation Knowledge 16 Jan 2022 · 0 repositories
-
That is a good looking car !: Visual Aspect based Sentiment Controlled Personalized Response Generation 16 Jan 2022 · 0 repositories
-
Tree Knowledge Distillation for Compressing Transformer-Based Language Models 16 Jan 2022 · 0 repositories
-
Uncovering Surprising Event Boundaries in Narratives 16 Jan 2022 · 0 repositories
-
Understand before Answer: Improve Temporal Reading Comprehension via Precise Question Understanding 16 Jan 2022 · 0 repositories
-
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models 16 Jan 2022 · 1 repository · arXiv:2201.05966Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
VEE-BERT: Accelerating BERT Inference for Named Entity Recognition via Vote Early Exiting 16 Jan 2022 · 0 repositories
-
Video Transformers: A Survey 16 Jan 2022 · 0 repositories · arXiv:2201.05991
-
WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation 16 Jan 2022 · 1 repository · arXiv:2201.05955
-
What do tokens know about their characters and how do they know it? 16 Jan 2022 · 1 repository
-
What Role Does BERT Play in the Neural Machine Translation Encoder? 16 Jan 2022 · 0 repositories
-
When a sentence does not introduce a discourse entity, Transformer-based models still often refer to it 16 Jan 2022 · 0 repositories
-
Why Does Surprisal From Smaller GPT-2 Models Provide Better Fit to Human Reading Times? 16 Jan 2022 · 0 repositories
-
Automatic Correction of Syntactic Dependency Annotation Differences 15 Jan 2022 · 0 repositories · arXiv:2201.05891
-
Automatic Lexical Simplification for Turkish 15 Jan 2022 · 0 repositories · arXiv:2201.05878
-
Domain Adaptation via Bidirectional Cross-Attention Transformer 15 Jan 2022 · 0 repositories · arXiv:2201.05887
-
Kformer: Knowledge Injection in Transformer Feed-Forward Layers 15 Jan 2022 · 1 repository · arXiv:2201.05742Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Machine Learning for Food Review and Recommendation 15 Jan 2022 · 0 repositories · arXiv:2201.10978
-
ViTBIS: Vision Transformer for Biomedical Image Segmentation 15 Jan 2022 · 0 repositories · arXiv:2201.05920
-
Applying a Generic Sequence-to-Sequence Model for Simple and Effective Keyphrase Generation 14 Jan 2022 · 0 repositories · arXiv:2201.05302
-
Attention over Self-attention:Intention-aware Re-ranking with Dynamic Transformer Encoders for Recommendation 14 Jan 2022 · 0 repositories · arXiv:2201.05333
-
CommonsenseQA 2.0: Exposing the Limits of AI through Gamification 14 Jan 2022 · 0 repositories · arXiv:2201.05320
-
Polarity and Subjectivity Detection with Multitask Learning and BERT Embedding 14 Jan 2022 · 0 repositories · arXiv:2201.05363
-
Accurate identification of bacteriophages from metagenomic data using Transformer 13 Jan 2022 · 1 repository · arXiv:2201.04778
-
Assemble Foundation Models for Automatic Code Summarization 13 Jan 2022 · 1 repository · arXiv:2201.05222
-
Hand-Object Interaction Reasoning 13 Jan 2022 · 0 repositories · arXiv:2201.04906
-
Knowledge Graph Augmented Network Towards Multiview Representation Learning for Aspect-based Sentiment Analysis 13 Jan 2022 · 1 repository · arXiv:2201.04831
-
Multi-task Pre-training Language Model for Semantic Network Completion 13 Jan 2022 · 1 repository · arXiv:2201.04843
-
Technical Report for ICCV 2021 Challenge SSLAD-Track3B: Transformers Are Better Continual Learners 13 Jan 2022 · 0 repositories · arXiv:2201.04924
-
Towards Automated Error Analysis: Learning to Characterize Errors 13 Jan 2022 · 0 repositories · arXiv:2201.05017
-
TransVOD: End-to-End Video Object Detection with Spatial-Temporal Transformers 13 Jan 2022 · 3 repositories · arXiv:2201.05047Syntology official (archive's flag): 3 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Chaos and order in event-triggered control 12 Jan 2022 · 0 repositories · arXiv:2201.04462
-
Diagnosing BERT with Retrieval Heuristics 12 Jan 2022 · 1 repository · arXiv:2201.04458
-
Generative Adversarial Network for Text-to-Face Synthesis and Manipulation with Pretrained BERT Model 12 Jan 2022 · 0 repositories
-
PromptBERT: Improving BERT Sentence Embeddings with Prompts 12 Jan 2022 · 1 repository · arXiv:2201.04337Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; the one sample that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
A Feature Extraction based Model for Hate Speech Identification 11 Jan 2022 · 0 repositories · arXiv:2201.04227
-
Explaining Predictive Uncertainty by Looking Back at Model Explanations 11 Jan 2022 · 0 repositories · arXiv:2201.03742
-
HyperTransformer: Model Generation for Supervised and Semi-Supervised Few-Shot Learning 11 Jan 2022 · 2 repositories · arXiv:2201.04182
-
Model-less Robust Voltage Control in Active Distribution Networks using Sensitivity Coefficients Estimated from Measurements 11 Jan 2022 · 0 repositories · arXiv:2201.04192
-
Pyramid Fusion Transformer for Semantic Segmentation 11 Jan 2022 · 0 repositories · arXiv:2201.04019
-
Quantifying Robustness to Adversarial Word Substitutions 11 Jan 2022 · 0 repositories · arXiv:2201.03829
-
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild 11 Jan 2022 · 0 repositories · arXiv:2201.03857
-
Uni-EDEN: Universal Encoder-Decoder Network by Multi-Granular Vision-Language Pre-training 11 Jan 2022 · 0 repositories · arXiv:2201.04026
-
BERT for Sentiment Analysis: Pre-trained and Fine-Tuned Alternatives 10 Jan 2022 · 2 repositories · arXiv:2201.03382
-
Black-Box Tuning for Language-Model-as-a-Service 10 Jan 2022 · 2 repositories · arXiv:2201.03514Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Handwriting recognition and automatic scoring for descriptive answers in Japanese language tests 10 Jan 2022 · 0 repositories · arXiv:2201.03215
-
Local Information Assisted Attention-free Decoder for Audio Captioning 10 Jan 2022 · 1 repository · arXiv:2201.03217
-
SCROLLS: Standardized CompaRison Over Long Language Sequences 10 Jan 2022 · 2 repositories · arXiv:2201.03533Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Swin Transformer for Fast MRI 10 Jan 2022 · 2 repositories · arXiv:2201.03230Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Swin Transformer coupling CNNs Makes Strong Contextual Encoders for VHR Image Road Extraction 10 Jan 2022 · 0 repositories · arXiv:2201.03178
-
Latency Adjustable Transformer Encoder for Language Understanding 10 Jan 2022 · 0 repositories · arXiv:2201.03327
-
Spatio-Temporal Tuples Transformer for Skeleton-Based Action Recognition 8 Jan 2022 · 1 repository · arXiv:2201.02849Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset 7 Jan 2022 · 1 repository · arXiv:2201.02419
-
Imagined versus Remembered Stories: Quantifying Differences in Narrative Flow 7 Jan 2022 · 0 repositories · arXiv:2201.02662
-
Learning Target-aware Representation for Visual Tracking via Informative Interactions 7 Jan 2022 · 0 repositories · arXiv:2201.02526
-
Compact Bidirectional Transformer for Image Captioning 6 Jan 2022 · 1 repository · arXiv:2201.01984
-
Flow-Guided Sparse Transformer for Video Deblurring 6 Jan 2022 · 1 repository · arXiv:2201.01893Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Self-Training Vision Language BERTs with a Unified Conditional Model 6 Jan 2022 · 0 repositories · arXiv:2201.02010
-
TransVPR: Transformer-based place recognition with multi-level attention aggregation 6 Jan 2022 · 0 repositories · arXiv:2201.02001
-
Formal Analysis of Art: Proxy Learning of Visual Concepts from Style Through Language Models 5 Jan 2022 · 0 repositories · arXiv:2201.01819
-
Lawin Transformer: Improving Semantic Segmentation Transformer with Multi-Scale Representations via Large Window Attention 5 Jan 2022 · 3 repositories · arXiv:2201.01615
-
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction 5 Jan 2022 · 2 repositories · arXiv:2201.02184
-
Comparison of biomedical relationship extraction methods and models for knowledge graph creation 5 Jan 2022 · 0 repositories · arXiv:2201.01647
-
Sparse-Dyn: Sparse Dynamic Graph Multi-representation Learning via Event-based Sparse Temporal Attention Network 4 Jan 2022 · 0 repositories · arXiv:2201.01384
-
PyramidTNT: Improved Transformer-in-Transformer Baselines with Pyramid Architecture 4 Jan 2022 · 1 repository · arXiv:2201.00978
-
Short Range Correlation Transformer for Occluded Person Re-Identification 4 Jan 2022 · 0 repositories · arXiv:2201.01090
-
Sign Pose-Based Transformer for Word-Level Sign Language Recognition 4 Jan 2022 · 1 repository
-
Submix: Practical Private Prediction for Large-Scale Language Models 4 Jan 2022 · 0 repositories · arXiv:2201.00971
-
Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images 4 Jan 2022 · 3 repositories · arXiv:2201.01266
-
An Adversarial Benchmark for Fake News Detection Models 3 Jan 2022 · 1 repository · arXiv:2201.00912
-
CaFT: Clustering and Filter on Tokens of Transformer for Weakly Supervised Object Localization 3 Jan 2022 · 0 repositories · arXiv:2201.00475
-
D-Former: A U-shaped Dilated Transformer for 3D Medical Image Segmentation 3 Jan 2022 · 1 repository · arXiv:2201.00462
-
Language as Queries for Referring Video Object Segmentation 3 Jan 2022 · 1 repository · arXiv:2201.00487Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization Space 3 Jan 2022 · 1 repository · arXiv:2201.00814Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Vision Transformer with Deformable Attention 3 Jan 2022 · 2 repositories · arXiv:2201.00520Syntology official (archive's flag): 2 ran · 8 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Which Student is Best? A Comprehensive Knowledge Distillation Exam for Task-Specific BERT Models 3 Jan 2022 · 0 repositories · arXiv:2201.00558
-
Detail-Preserving Transformer for Light Field Image Super-Resolution 2 Jan 2022 · 1 repository · arXiv:2201.00346Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 4 where Syntology's instrument failed) · 11 unverified (of 21 harvested samples) · 21 pointer-only (licence)
-
Informed Multi-context Entity Alignment 2 Jan 2022 · 1 repository · arXiv:2201.00304
-
On Sensitivity of Deep Learning Based Text Classification Algorithms to Practical Input Perturbations 2 Jan 2022 · 0 repositories · arXiv:2201.00318
-
Splicing ViT Features for Semantic Appearance Transfer 2 Jan 2022 · 1 repository · arXiv:2201.00424Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
A Brand New Dance Partner: Music-Conditioned Pluralistic Dancing Controlled by Multiple Dance Genres 1 Jan 2022 · 1 repository
-
A Graph Matching Perspective With Transformers on Video Instance Segmentation 1 Jan 2022 · 0 repositories
-
CADTransformer: Panoptic Symbol Spotting Transformer for CAD Drawings 1 Jan 2022 · 1 repository
-
Chitransformer: Towards Reliable Stereo From Cues 1 Jan 2022 · 1 repository
-
Continual Learning With Lifelong Vision Transformer 1 Jan 2022 · 0 repositories
-
Continual Stereo Matching of Continuous Driving Scenes With Growing Architecture 1 Jan 2022 · 1 repository
-
DESTR: Object Detection With Split Transformer 1 Jan 2022 · 0 repositories
-
DLFormer: Discrete Latent Transformer for Video Inpainting 1 Jan 2022 · 0 repositories
-
Dynamic Scene Graph Generation via Anticipatory Pre-Training 1 Jan 2022 · 0 repositories
-
Expanding Large Pre-Trained Unimodal Models With Multimodal Information Injection for Image-Text Multimodal Classification 1 Jan 2022 · 0 repositories
-
HiVT: Hierarchical Vector Transformer for Multi-Agent Motion Prediction 1 Jan 2022 · 2 repositories
-
Image Dehazing Transformer With Transmission-Aware 3D Position Embedding 1 Jan 2022 · 2 repositories
-
Instance Segmentation With Mask-Supervised Polygonal Boundary Transformers 1 Jan 2022 · 0 repositories
-
KNN Local Attention for Image Restoration 1 Jan 2022 · 0 repositories
-
Learning Transferable Human-Object Interaction Detector With Natural Language Supervision 1 Jan 2022 · 1 repository
-
LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object Detection 1 Jan 2022 · 0 repositories
-
Likert Scoring With Grade Decoupling for Long-Term Action Assessment 1 Jan 2022 · 0 repositories
-
LTP: Lane-Based Trajectory Prediction for Autonomous Driving 1 Jan 2022 · 0 repositories
-
M3T: Three-Dimensional Medical Image Classifier Using Multi-Plane and Multi-Slice Transformer 1 Jan 2022 · 1 repository
-
Mr.BiQ: Post-Training Non-Uniform Quantization Based on Minimizing the Reconstruction Error 1 Jan 2022 · 0 repositories