Methods › General › Attention Modules › Multi-Head Attention › Papers, page 173
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 173 of 249: papers 17,201 to 17,300 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Mutual Distillation Learning Network for Trajectory-User Linking 8 May 2022 · 1 repository · arXiv:2205.03773Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
On the Use of BERT for Automated Essay Scoring: Joint Learning of Multi-Scale Essay Representation 8 May 2022 · 1 repository · arXiv:2205.03835
-
AKI-BERT: a Pre-trained Clinical Language Model for Early Prediction of Acute Kidney Injury 7 May 2022 · 1 repository · arXiv:2205.03695
-
EmotionFlow: Capture the Dialogue Level Emotion Transitions 7 May 2022 · 1 repository
-
Improving Downstream Task Performance by Treating Numbers as Entities 7 May 2022 · 0 repositories · arXiv:2205.03559
-
Vector Representations of Idioms in Conversational Systems 7 May 2022 · 0 repositories · arXiv:2205.03666
-
A Data Cartography based MixUp for Pre-trained Language Models 6 May 2022 · 1 repository · arXiv:2205.03403
-
Explaining the Effectiveness of Multi-Task Learning for Efficient Knowledge Extraction from Spine MRI Reports 6 May 2022 · 0 repositories · arXiv:2205.02979
-
Fake News Detection with Heterogeneous Transformer 6 May 2022 · 1 repository · arXiv:2205.03100
-
RCMNet: A deep learning model assists CAR-T therapy for leukemia 6 May 2022 · 0 repositories · arXiv:2205.04230
-
Stock Price Prediction Based on Natural Language Processing 6 May 2022 · 2 repositories
-
The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning 6 May 2022 · 1 repository · arXiv:2205.03401
-
Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment 6 May 2022 · 1 repository · arXiv:2205.03432
-
When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it 6 May 2022 · 1 repository · arXiv:2205.03472
-
A Deep Learning Approach to Dst Index Prediction 5 May 2022 · 0 repositories · arXiv:2205.02447
-
BORT: Back and Denoising Reconstruction for End-to-End Task-Oriented Dialog 5 May 2022 · 1 repository · arXiv:2205.02471
-
Declaration-based Prompt Tuning for Visual Question Answering 5 May 2022 · 1 repository · arXiv:2205.02456
-
Exploiting Global and Local Hierarchies for Hierarchical Text Classification 5 May 2022 · 1 repository · arXiv:2205.02613
-
Identifying Cause-and-Effect Relationships of Manufacturing Errors using Sequence-to-Sequence Learning 5 May 2022 · 0 repositories · arXiv:2205.02827
-
RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced Labour 5 May 2022 · 0 repositories · arXiv:2205.02684
-
Scene Graph Expansion for Semantics-Guided Image Outpainting 5 May 2022 · 0 repositories · arXiv:2205.02958
-
Hyperbolic Relevance Matching for Neural Keyphrase Extraction 4 May 2022 · 1 repository · arXiv:2205.02047
-
Improving Multi-Document Summarization through Referenced Flexible Extraction with Credit-Awareness 4 May 2022 · 1 repository · arXiv:2205.01889
-
Knowledge Distillation of Russian Language Models with Reduction of Vocabulary 4 May 2022 · 1 repository · arXiv:2205.02340
-
Provably Confidential Language Modelling 4 May 2022 · 1 repository · arXiv:2205.01863Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Using virtual edges to extract keywords from texts modeled as complex networks 4 May 2022 · 0 repositories · arXiv:2205.02172
-
Better plain ViT baselines for ImageNet-1k 3 May 2022 · 7 repositories · arXiv:2205.01580Syntology official: no sample here; runs from other or unrecorded repositories · 23 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 5 honoured, 1 violated, 14 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 24 harvested samples) · 4 pointer-only (licence)
-
Contrastive Learning for Prompt-Based Few-Shot Language Learners 3 May 2022 · 1 repository · arXiv:2205.01308Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 6 pointer-only (licence)
-
MTTrans: Cross-Domain Object Detection with Mean-Teacher Transformer 3 May 2022 · 1 repository · arXiv:2205.01643
-
Efficient Fine-Tuning of BERT Models on the Edge 3 May 2022 · 0 repositories · arXiv:2205.01541
-
Explain and Conquer: Personalised Text-based Reviews to Achieve Transparency 3 May 2022 · 0 repositories · arXiv:2205.01759
-
Finding patterns in Knowledge Attribution for Transformers 3 May 2022 · 0 repositories · arXiv:2205.01366
-
Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation 3 May 2022 · 0 repositories · arXiv:2205.01546
-
Mixed-effects transformers for hierarchical adaptation 3 May 2022 · 1 repository · arXiv:2205.01749
-
Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity? 3 May 2022 · 0 repositories · arXiv:2205.01404
-
Predicting Issue Types with seBERT 3 May 2022 · 1 repository · arXiv:2205.01335
-
SemAttack: Natural Textual Attacks via Different Semantic Spaces 3 May 2022 · 1 repository · arXiv:2205.01287Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Synthesized Speech Detection Using Convolutional Transformer-Based Spectrogram Analysis 3 May 2022 · 0 repositories · arXiv:2205.01800
-
Textual Entailment for Event Argument Extraction: Zero- and Few-Shot with Multi-Source Learning 3 May 2022 · 1 repository · arXiv:2205.01376
-
BERTops: Studying BERT Representations under a Topological Lens 2 May 2022 · 1 repository · arXiv:2205.00953
-
CenterCLIP: Token Clustering for Efficient Text-Video Retrieval 2 May 2022 · 1 repository · arXiv:2205.00823Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Entity-aware Transformers for Entity Search 2 May 2022 · 1 repository · arXiv:2205.00820
-
Improving Students' Academic Performance with AI and Semantic Technologies 2 May 2022 · 1 repository · arXiv:2206.03213
-
Logiformer: A Two-Branch Graph Transformer Network for Interpretable Logical Reasoning 2 May 2022 · 1 repository · arXiv:2205.00731
-
Multi-Task Text Classification using Graph Convolutional Networks for Large-Scale Low Resource Language 2 May 2022 · 1 repository · arXiv:2205.01204
-
OPT: Open Pre-trained Transformer Language Models 2 May 2022 · 11 repositories · arXiv:2205.01068Syntology official (archive's flag): 9 ran · 14 ran (of which 2 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 17 pointer-only (licence)
-
Teaching BERT to Wait: Balancing Accuracy and Latency for Streaming Disfluency Detection 2 May 2022 · 0 repositories · arXiv:2205.00620
-
Aircraft engine remaining useful life estimation via a double attention-based data-driven architecture 1 May 2022 · 0 repositories
-
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis 1 May 2022 · 0 repositories
-
Detecting COVID-19 Conspiracy Theories with Transformers and TF-IDF 1 May 2022 · 0 repositories · arXiv:2205.00377
-
Reinforced Swin-Convs Transformer for Underwater Image Enhancement 1 May 2022 · 1 repository · arXiv:2205.00434Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 8 harvested samples)
-
Coarse-to-Fine Video Denoising with Dual-Stage Spatial-Channel Transformer 30 Apr 2022 · 0 repositories · arXiv:2205.00214
-
HDGT: Heterogeneous Driving Graph Transformer for Multi-Agent Trajectory Prediction via Scene Encoding 30 Apr 2022 · 2 repositories · arXiv:2205.09753
-
StorSeismic: A new paradigm in deep learning for seismic processing 30 Apr 2022 · 1 repository · arXiv:2205.00222
-
Unsupervised Contrastive Learning based Transformer for Lung Nodule Detection 30 Apr 2022 · 0 repositories · arXiv:2205.00122
-
ExaASC: A General Target-Based Stance Detection Corpus in Arabic Language 29 Apr 2022 · 1 repository · arXiv:2204.13979
-
Training Language Models with Language Feedback 29 Apr 2022 · 0 repositories · arXiv:2204.14146
-
QRelScore: Better Evaluating Generated Questions with Deeper Understanding of Context-aware Relevance 29 Apr 2022 · 0 repositories · arXiv:2204.13921
-
Two New Datasets for Italian-Language Abstractive Text Summarization 29 Apr 2022 · 1 repository
-
Controllable Image Captioning 28 Apr 2022 · 0 repositories · arXiv:2204.13324
-
Depth Estimation with Simplified Transformer 28 Apr 2022 · 0 repositories · arXiv:2204.13791
-
HiNER: A Large Hindi Named Entity Recognition Dataset 28 Apr 2022 · 1 repository · arXiv:2204.13743
-
Inferring Implicit Relations in Complex Questions with Language Models 28 Apr 2022 · 1 repository · arXiv:2204.13778
-
Lightweight Bimodal Network for Single-Image Super-Resolution via Symmetric CNN and Recursive Transformer 28 Apr 2022 · 1 repository · arXiv:2204.13286
-
On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model 28 Apr 2022 · 0 repositories · arXiv:2204.13509
-
One Model to Synthesize Them All: Multi-contrast Multi-scale Transformer for Missing Data Imputation 28 Apr 2022 · 0 repositories · arXiv:2204.13738
-
RobBERTje: a Distilled Dutch BERT Model 28 Apr 2022 · 0 repositories · arXiv:2204.13511
-
Symmetric Transformer-based Network for Unsupervised Image Registration 28 Apr 2022 · 1 repository · arXiv:2204.13575
-
Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing Modalities 28 Apr 2022 · 1 repository · arXiv:2204.13707
-
Tailor: A Prompt-Based Approach to Attribute-Based Controlled Text Generation 28 Apr 2022 · 0 repositories · arXiv:2204.13362
-
Transformers in Time-series Analysis: A Tutorial 28 Apr 2022 · 0 repositories · arXiv:2205.01138
-
UM6P-CS at SemEval-2022 Task 11: Enhancing Multilingual and Code-Mixed Complex Named Entity Recognition via Pseudo Labels using Multilingual Transformer 28 Apr 2022 · 0 repositories · arXiv:2204.13515
-
An End-to-End Dialogue Summarization System for Sales Calls 27 Apr 2022 · 0 repositories · arXiv:2204.12951
-
Better Query Graph Selection for Knowledge Base Question Answering 27 Apr 2022 · 0 repositories · arXiv:2204.12662
-
CATrans: Context and Affinity Transformer for Few-Shot Segmentation 27 Apr 2022 · 0 repositories · arXiv:2204.12817
-
Building Knowledge-Grounded Dialogue Systems with Graph-Based Semantic Modeling 27 Apr 2022 · 0 repositories · arXiv:2204.12681
-
Modern Baselines for SPARQL Semantic Parsing 27 Apr 2022 · 1 repository · arXiv:2204.12793
-
RigoBERTa: A State-of-the-Art Language Model For Spanish 27 Apr 2022 · 0 repositories · arXiv:2205.10233
-
SkillSpan: Hard and Soft Skill Extraction from English Job Postings 27 Apr 2022 · 1 repository · arXiv:2204.12811
-
BiTimeBERT: Extending Pre-Trained Language Representations with Bi-Temporal Information 27 Apr 2022 · 0 repositories · arXiv:2204.13032
-
Ultra Fast Speech Separation Model with Teacher Student Learning 27 Apr 2022 · 0 repositories · arXiv:2204.12777
-
A survey on attention mechanisms for medical applications: are we moving towards better algorithms? 26 Apr 2022 · 1 repository · arXiv:2204.12406
-
Adaptive Split-Fusion Transformer 26 Apr 2022 · 1 repository · arXiv:2204.12196
-
MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval 26 Apr 2022 · 1 repository · arXiv:2204.12408
-
PLOD: An Abbreviation Detection Dataset for Scientific Documents 26 Apr 2022 · 1 repository · arXiv:2204.12061
-
Pretraining Chinese BERT for Detecting Word Insertion and Deletion Errors 26 Apr 2022 · 0 repositories · arXiv:2204.12052
-
Thompson Sampling for Bandit Learning in Matching Markets 26 Apr 2022 · 1 repository · arXiv:2204.12048
-
ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation 26 Apr 2022 · 6 repositories · arXiv:2204.12484Syntology community repositories only · 18 ran (of which 8 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 6 where Syntology's instrument failed) · 13 unverified (of 31 harvested samples) · 6 pointer-only (licence)
-
Crystal Transformer: Self-learning neural language model for Generative and Tinkering Design of Materials 25 Apr 2022 · 0 repositories · arXiv:2204.11953
-
Evaluating Interpolation and Extrapolation Performance of Neural Retrieval Models 25 Apr 2022 · 1 repository · arXiv:2204.11447
-
Groupwise Query Performance Prediction with BERT 25 Apr 2022 · 1 repository · arXiv:2204.11489
-
OCFormer: One-Class Transformer Network for Image Classification 25 Apr 2022 · 0 repositories · arXiv:2204.11449
-
Performer: A Novel PPG-to-ECG Reconstruction Transformer for a Digital Biomarker of Cardiovascular Disease Detection 25 Apr 2022 · 0 repositories · arXiv:2204.11795
-
Predicting Real-time Scientific Experiments Using Transformer models and Reinforcement Learning 25 Apr 2022 · 1 repository · arXiv:2204.11718
-
SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images 25 Apr 2022 · 1 repository · arXiv:2204.11436
-
Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation 24 Apr 2022 · 1 repository · arXiv:2204.11320
-
Faster Learned Sparse Retrieval with Guided Traversal 24 Apr 2022 · 1 repository · arXiv:2204.11314
-
Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training 24 Apr 2022 · 1 repository · arXiv:2204.11218
-
RelViT: Concept-guided Vision Transformer for Visual Relational Reasoning 24 Apr 2022 · 1 repository · arXiv:2204.11167Syntology official (archive's flag): 6 ran · 6 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Local Gaussian process extrapolation for BART models with applications to causal inference 23 Apr 2022 · 0 repositories · arXiv:2204.10963