Methods › General › Attention Modules › Multi-Head Attention › Papers, page 190
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 190 of 249: papers 18,901 to 19,000 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Effective Cross-Utterance Language Modeling for Conversational Speech Recognition 5 Nov 2021 · 0 repositories · arXiv:2111.03333
-
Hepatic vessel segmentation based on 3D swin-transformer with inductive biased multi-head self-attention 5 Nov 2021 · 0 repositories · arXiv:2111.03368
-
IBERT: Idiom Cloze-style reading comprehension with Attention 5 Nov 2021 · 0 repositories · arXiv:2112.02994
-
Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers 5 Nov 2021 · 0 repositories · arXiv:2111.03481
-
Oracle Teacher: Leveraging Target Information for Better Knowledge Distillation of CTC Models 5 Nov 2021 · 0 repositories · arXiv:2111.03664
-
Sexism Identification in Tweets and Gabs using Deep Neural Networks 5 Nov 2021 · 0 repositories · arXiv:2111.03612
-
A text autoencoder from transformer for fast encoding language representation 4 Nov 2021 · 0 repositories · arXiv:2111.02844
-
An Empirical Study of the Effectiveness of an Ensemble of Stand-alone Sentiment Detection Tools for Software Engineering Datasets 4 Nov 2021 · 1 repository · arXiv:2111.03196
-
Benchmarking Multimodal AutoML for Tabular Data with Text Fields 4 Nov 2021 · 2 repositories · arXiv:2111.02705Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Conformal prediction for text infilling and part-of-speech prediction 4 Nov 2021 · 1 repository · arXiv:2111.02592
-
Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports 4 Nov 2021 · 1 repository · arXiv:2111.03452Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
MT3: Multi-Task Multitrack Music Transcription 4 Nov 2021 · 3 repositories · arXiv:2111.03017Syntology official: harvested, nothing ran · 0 ran · 6 unverified (of 6 harvested samples)
-
Multi-Airport Delay Prediction with Transformers 4 Nov 2021 · 0 repositories · arXiv:2111.04494
-
An Empirical Study of Training End-to-End Vision-and-Language Transformers 3 Nov 2021 · 3 repositories · arXiv:2111.02387Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
An Explanation of In-context Learning as Implicit Bayesian Inference 3 Nov 2021 · 1 repository · arXiv:2111.02080Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
BERT-DRE: BERT with Deep Recursive Encoder for Natural Language Sentence Matching 3 Nov 2021 · 0 repositories · arXiv:2111.02188
-
ProSTformer: Pre-trained Progressive Space-Time Self-attention Model for Traffic Flow Forecasting 3 Nov 2021 · 0 repositories · arXiv:2111.03459
-
TheEyeCorpus: Experiments in Reducing NLP Bias and Identifiability for Large LMs 3 Nov 2021 · 0 repositories
-
TranSMS: Transformers for Super-Resolution Calibration in Magnetic Particle Imaging 3 Nov 2021 · 2 repositories · arXiv:2111.02163
-
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts 3 Nov 2021 · 2 repositories · arXiv:2111.02358
-
Can Vision Transformers Perform Convolution? 2 Nov 2021 · 0 repositories · arXiv:2111.01353
-
Detection of Hate Speech using BERT and Hate Speech Word Embedding with Deep Model 2 Nov 2021 · 0 repositories · arXiv:2111.01515
-
Explaining Documents' Relevance to Search Queries 2 Nov 2021 · 0 repositories · arXiv:2111.01314
-
Federated Split Vision Transformer for COVID-19 CXR Diagnosis using Task-Agnostic Training 2 Nov 2021 · 0 repositories · arXiv:2111.01338
-
Relational Self-Attention: What's Missing in Attention for Video Understanding 2 Nov 2021 · 1 repository · arXiv:2111.01673Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Sentence encoding for Dialogue Act classification 2 Nov 2021 · 1 repository
-
UQuAD1.0: Development of an Urdu Question Answering Training Data for Machine Reading Comprehension 2 Nov 2021 · 0 repositories · arXiv:2111.01543
-
Accounting for Dependencies in Deep Learning Based Multiple Instance Learning for Whole Slide Imaging 1 Nov 2021 · 0 repositories · arXiv:2111.01556
-
Arch-Net: Model Distillation for Architecture Agnostic Model Deployment 1 Nov 2021 · 1 repository · arXiv:2111.01135
-
Comparative Study of Long Document Classification 1 Nov 2021 · 0 repositories · arXiv:2111.00702
-
Cross-lingual Hate Speech Detection using Transformer Models 1 Nov 2021 · 0 repositories · arXiv:2111.00981
-
Identifying causal relations in tweets using deep learning: Use case on diabetes-related tweets from 2017-2021 1 Nov 2021 · 1 repository · arXiv:2111.01225
-
Learning To Generate Piano Music With Sustain Pedals 1 Nov 2021 · 1 repository · arXiv:2111.01216
-
MAPLE – MAsking words to generate blackout Poetry using sequence-to-sequence LEarning 1 Nov 2021 · 1 repository
-
Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey 1 Nov 2021 · 0 repositories · arXiv:2111.01243
-
VSEC: Transformer-based Model for Vietnamese Spelling Correction 1 Nov 2021 · 0 repositories · arXiv:2111.00640
-
Wino-X: Multilingual Winograd Schemas for Commonsense Reasoning and Coreference Resolution 1 Nov 2021 · 1 repository
-
FinEAS: Financial Embedding Analysis of Sentiment 31 Oct 2021 · 1 repository · arXiv:2111.00526
-
Backdoor Pre-trained Models Can Transfer to All 30 Oct 2021 · 1 repository · arXiv:2111.00197
-
Cross-Modality Fusion Transformer for Multispectral Object Detection 30 Oct 2021 · 1 repository · arXiv:2111.00273
-
DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language Models 30 Oct 2021 · 1 repository · arXiv:2111.00160Syntology official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 10 pointer-only (licence)
-
Magic Pyramid: Accelerating Inference with Early Exiting and Token Pruning 30 Oct 2021 · 0 repositories · arXiv:2111.00230
-
PatchFormer: An Efficient Point Transformer with Patch Attention 30 Oct 2021 · 0 repositories · arXiv:2111.00207
-
Amendable Generation for Dialogue State Tracking 29 Oct 2021 · 1 repository · arXiv:2110.15659
-
Delayed Propagation Transformer: A Universal Computation Engine towards Practical Control in Cyber-Physical Systems 29 Oct 2021 · 1 repository · arXiv:2110.15926Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
REBEL: Relation Extraction By End-to-end Language generation 29 Oct 2021 · 1 repository
-
Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR Parsing 29 Oct 2021 · 1 repository · arXiv:2110.15534
-
The Golden Rule as a Heuristic to Measure the Fairness of Texts Using Machine Learning 29 Oct 2021 · 0 repositories · arXiv:2111.00107
-
ICDM 2020 Knowledge Graph Contest: Consumer Event-Cause Extraction 28 Oct 2021 · 0 repositories · arXiv:2110.15722
-
A Sequence to Sequence Model for Extracting Multiple Product Name Entities from Dialog 28 Oct 2021 · 0 repositories · arXiv:2110.14843
-
Blending Anti-Aliasing into Vision Transformer 28 Oct 2021 · 0 repositories · arXiv:2110.15156
-
Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework 28 Oct 2021 · 1 repository · arXiv:2110.15317
-
Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training 28 Oct 2021 · 1 repository · arXiv:2110.14883
-
Dispensed Transformer Network for Unsupervised Domain Adaptation 28 Oct 2021 · 0 repositories · arXiv:2110.14944
-
NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMM 28 Oct 2021 · 0 repositories · arXiv:2110.15766
-
Pruning Attention Heads of Transformer Models Using A* Search: A Novel Approach to Compress Big NLP Architectures 28 Oct 2021 · 0 repositories · arXiv:2110.15225
-
Scatterbrain: Unifying Sparse and Low-rank Attention Approximation 28 Oct 2021 · 1 repository · arXiv:2110.15343Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Hyper-Representations: Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction 28 Oct 2021 · 1 repository · arXiv:2110.15288Syntology official (archive's flag): 9 ran · 9 ran (of which 6 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Semi-Siamese Bi-encoder Neural Ranking Model Using Lightweight Fine-Tuning 28 Oct 2021 · 1 repository · arXiv:2110.14943
-
Anomaly-Injected Deep Support Vector Data Description for Text Outlier Detection 27 Oct 2021 · 0 repositories · arXiv:2110.14729
-
Ask me in your own words: paraphrasing for multitask question answering 27 Oct 2021 · 1 repository
-
Detecting Dementia from Speech and Transcripts using Transformers 27 Oct 2021 · 0 repositories · arXiv:2110.14769
-
Discovering Non-monotonic Autoregressive Orderings with Variational Inference 27 Oct 2021 · 1 repository · arXiv:2110.15797Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 7 harvested samples)
-
Transfer learning with causal counterfactual reasoning in Decision Transformers 27 Oct 2021 · 0 repositories · arXiv:2110.14355
-
Transformers Generalize DeepSets and Can be Extended to Graphs and Hypergraphs 27 Oct 2021 · 2 repositories · arXiv:2110.14416Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Vision Transformer for Classification of Breast Ultrasound Images 27 Oct 2021 · 0 repositories · arXiv:2110.14731
-
Can't Fool Me: Adversarially Robust Transformer for Video Understanding 26 Oct 2021 · 0 repositories · arXiv:2110.13950
-
CLAUSEREC: A Clause Recommendation Framework for AI-aided Contract Authoring 26 Oct 2021 · 0 repositories · arXiv:2110.15794
-
Geometric Transformer for End-to-End Molecule Properties Prediction 26 Oct 2021 · 1 repository · arXiv:2110.13721Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Hierarchical Transformers Are More Efficient Language Models 26 Oct 2021 · 3 repositories · arXiv:2110.13711Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 2 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
Leveraging Local Temporal Information for Multimodal Scene Classification 26 Oct 2021 · 0 repositories · arXiv:2110.13992
-
Post-processing for Individual Fairness 26 Oct 2021 · 1 repository · arXiv:2110.13796Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
s2s-ft: Fine-Tuning Pretrained Transformer Encoders for Sequence-to-Sequence Learning 26 Oct 2021 · 1 repository · arXiv:2110.13640
-
TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation 26 Oct 2021 · 1 repository · arXiv:2110.13412Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing 26 Oct 2021 · 9 repositories · arXiv:2110.13900
-
Actions Speak Louder than Listening: Evaluating Music Style Transfer based on Editing Experience 25 Oct 2021 · 1 repository · arXiv:2110.12855
-
DocTr: Document Image Transformer for Geometric Unwarping and Illumination Correction 25 Oct 2021 · 2 repositories · arXiv:2110.12942
-
Fine-tuning of Pre-trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi 25 Oct 2021 · 1 repository · arXiv:2110.12687
-
Generating artificial texts as substitution or complement of training data 25 Oct 2021 · 0 repositories · arXiv:2110.13016
-
Gophormer: Ego-Graph Transformer for Node Classification 25 Oct 2021 · 0 repositories · arXiv:2110.13094
-
History Aware Multimodal Transformer for Vision-and-Language Navigation 25 Oct 2021 · 1 repository · arXiv:2110.13309Syntology 7 ran (of which 3 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 15 harvested samples)
-
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning 25 Oct 2021 · 1 repository · arXiv:2110.13214Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
MVT: Multi-view Vision Transformer for 3D Object Recognition 25 Oct 2021 · 2 repositories · arXiv:2110.13083Syntology official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Paradigm Shift in Language Modeling: Revisiting CNN for Modeling Sanskrit Originated Bengali and Hindi Language 25 Oct 2021 · 0 repositories · arXiv:2110.13032
-
The Nuts and Bolts of Adopting Transformer in GANs 25 Oct 2021 · 0 repositories · arXiv:2110.13107
-
CvT-ASSD: Convolutional vision-Transformer Based Attentive Single Shot MultiBox Detector 24 Oct 2021 · 1 repository · arXiv:2110.12364
-
Hate and Offensive Speech Detection in Hindi and Marathi 23 Oct 2021 · 0 repositories · arXiv:2110.12200
-
Bandits with Dynamic Arm-acquisition Costs 23 Oct 2021 · 0 repositories · arXiv:2110.12118
-
Double Trouble: How to not explain a text classifier's decisions using counterfactuals synthesized by masked language models? 22 Oct 2021 · 1 repository · arXiv:2110.11929
-
Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering 22 Oct 2021 · 1 repository · arXiv:2110.11592
-
Grafting Transformer on Automatically Designed Convolutional Neural Network for Hyperspectral Image Classification 21 Oct 2021 · 1 repository · arXiv:2110.11084
-
CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP 21 Oct 2021 · 1 repository · arXiv:2110.11316Syntology official (archive's flag): 4 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Fast Model Editing at Scale 21 Oct 2021 · 3 repositories · arXiv:2110.11309Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Modeling Performance in Open-Domain Dialogue with PARADISE 21 Oct 2021 · 0 repositories · arXiv:2110.11164
-
Transformer Acceleration with Dynamic Sparse Attention 21 Oct 2021 · 0 repositories · arXiv:2110.11299
-
Vis-TOP: Visual Transformer Overlay Processor 21 Oct 2021 · 0 repositories · arXiv:2110.10957
-
AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation 20 Oct 2021 · 0 repositories · arXiv:2110.10403
-
AniFormer: Data-driven 3D Animation with Transformer 20 Oct 2021 · 1 repository · arXiv:2110.10533
-
Continual Learning in Multilingual NMT via Language-Specific Embeddings 20 Oct 2021 · 0 repositories · arXiv:2110.10478
-
Distributionally Robust Classifiers in Sentiment Analysis 20 Oct 2021 · 1 repository · arXiv:2110.10372