Methods › General › Attention Modules › Multi-Head Attention › Papers, page 149
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 149 of 249: papers 14,801 to 14,900 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning 23 Jan 2023 · 1 repository · arXiv:2301.09626
-
Fully transformer-based biomarker prediction from colorectal cancer histology: a large-scale multicentric study 23 Jan 2023 · 2 repositories · arXiv:2301.09617
-
Injecting the BM25 Score as Text Improves BERT-Based Re-rankers 23 Jan 2023 · 1 repository · arXiv:2301.09728
-
ISTVT: Interpretable Spatial-Temporal Video Transformer for Deepfake Detection 23 Jan 2023 · 1 repository
-
Learning to View: Decision Transformers for Active Object Detection 23 Jan 2023 · 0 repositories · arXiv:2301.09544
-
Local Window Attention Transformer for Polarimetric SAR Image Classification 23 Jan 2023 · 1 repository
-
StockEmotions: Discover Investor Emotions for Financial Sentiment Analysis and Multivariate Time Series 23 Jan 2023 · 2 repositories · arXiv:2301.09279
-
Apples and Oranges? Assessing Image Quality over Content Recognition 22 Jan 2023 · 0 repositories · arXiv:2301.09190
-
Debiasing the Cloze Task in Sequential Recommendation with Bidirectional Transformers 22 Jan 2023 · 1 repository · arXiv:2301.09210
-
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien 21 Jan 2023 · 1 repository · arXiv:2301.08937
-
Exploring the Synergy Between Vision-Language Pretraining and ChatGPT for Artwork Captioning: A Preliminary Study 21 Jan 2023 · 1 repository
-
Slice Transformer and Self-supervised Learning for 6DoF Localization in 3D Point Cloud Maps 21 Jan 2023 · 0 repositories · arXiv:2301.08957
-
Stress Test for BERT and Deep Models: Predicting Words from Italian Poetry 21 Jan 2023 · 0 repositories · arXiv:2302.09303
-
SuperScaler: Supporting Flexible DNN Parallelization via a Unified Abstraction 21 Jan 2023 · 0 repositories · arXiv:2301.08984
-
Time-Conditioned Generative Modeling of Object-Centric Representations for Video Decomposition and Prediction 21 Jan 2023 · 1 repository · arXiv:2301.08951
-
Accelerating Multi-Agent Planning Using Graph Transformers with Bounded Suboptimality 20 Jan 2023 · 0 repositories · arXiv:2301.08451
-
Image Memorability Prediction with Vision Transformers 20 Jan 2023 · 0 repositories · arXiv:2301.08647
-
Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine 20 Jan 2023 · 1 repository · arXiv:2301.08745
-
Ontology Pre-training for Poison Prediction 20 Jan 2023 · 0 repositories · arXiv:2301.08577
-
Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions 20 Jan 2023 · 2 repositories · arXiv:2301.08810Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Which Features are Learned by CodeBert: An Empirical Study of the BERT-based Source Code Representation Learning 20 Jan 2023 · 0 repositories · arXiv:2301.08427
-
Batch Prompting: Efficient Inference with Large Language Model APIs 19 Jan 2023 · 2 repositories · arXiv:2301.08721
-
Diagnose Like a Pathologist: Transformer-Enabled Hierarchical Attention-Guided Multiple Instance Learning for Whole Slide Image Classification 19 Jan 2023 · 1 repository · arXiv:2301.08125
-
FE-TCM: Filter-Enhanced Transformer Click Model for Web Search 19 Jan 2023 · 0 repositories · arXiv:2301.07854
-
MedSegDiff-V2: Diffusion based Medical Image Segmentation with Transformer 19 Jan 2023 · 2 repositories · arXiv:2301.11798Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 0 violated, 10 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 21 harvested samples) · 10 pointer-only (licence)
-
Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing 18 Jan 2023 · 1 repository · arXiv:2301.07507
-
Cooperation Learning Enhanced Colonic Polyp Segmentation Based on Transformer-CNN Fusion 17 Jan 2023 · 0 repositories · arXiv:2301.06892
-
SAT: Size-Aware Transformer for 3D Point Cloud Semantic Segmentation 17 Jan 2023 · 0 repositories · arXiv:2301.06869
-
SwinDepth: Unsupervised Depth Estimation using Monocular Sequences via Swin Transformer and Densely Cascaded Network 17 Jan 2023 · 1 repository · arXiv:2301.06715
-
Tracing and Manipulating Intermediate Values in Neural Math Problem Solvers 17 Jan 2023 · 1 repository · arXiv:2301.06758
-
Transformer Based Implementation for Automatic Book Summarization 17 Jan 2023 · 0 repositories · arXiv:2301.07057
-
An Error-Guided Correction Model for Chinese Spelling Error Correction 16 Jan 2023 · 1 repository · arXiv:2301.06323
-
BayesSpeech: A Bayesian Transformer Network for Automatic Speech Recognition 16 Jan 2023 · 0 repositories · arXiv:2301.11276
-
Improving Target Speaker Extraction with Sparse LDA-transformed Speaker Embeddings 16 Jan 2023 · 0 repositories · arXiv:2301.06277
-
A Transformer-based Diffusion Probabilistic Model for Heart Rate and Blood Pressure Forecasting in Intensive Care Unit 16 Jan 2023 · 1 repository · arXiv:2301.06625
-
TEDB System Description to a Shared Task on Euphemism Detection 2022 16 Jan 2023 · 1 repository · arXiv:2301.06602
-
DSVT: Dynamic Sparse Voxel Transformer with Rotated Sets 15 Jan 2023 · 4 repositories · arXiv:2301.06051
-
Improving Noise Robustness for Spoken Content Retrieval using Semi-supervised ASR and N-best Transcripts for BERT-based Ranking Models 15 Jan 2023 · 0 repositories · arXiv:2301.06056
-
T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations 15 Jan 2023 · 1 repository · arXiv:2301.06052Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Salient Sign Detection In Safe Autonomous Driving: AI Which Reasons Over Full Visual Context 14 Jan 2023 · 0 repositories · arXiv:2301.05804
-
Automated speech- and text-based classification of neuropsychiatric conditions in a multidiagnostic setting 13 Jan 2023 · 0 repositories · arXiv:2301.06916
-
Efficient Activation Function Optimization through Surrogate Modeling 13 Jan 2023 · 2 repositories · arXiv:2301.05785Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Text to Point Cloud Localization with Relation-Enhanced Transformer 13 Jan 2023 · 0 repositories · arXiv:2301.05372
-
Adversarial Adaptation for French Named Entity Recognition 12 Jan 2023 · 1 repository · arXiv:2301.05220
-
ViTs for SITS: Vision Transformers for Satellite Image Time Series 12 Jan 2023 · 3 repositories · arXiv:2301.04944Syntology official (archive's flag): 3 ran · 11 ran (of which 3 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
AdaPoinTr: Diverse Point Cloud Completion with Adaptive Geometry-Aware Transformers 11 Jan 2023 · 1 repository · arXiv:2301.04545Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Anomalies, Representations, and Self-Supervision 11 Jan 2023 · 0 repositories · arXiv:2301.04660
-
Super-resolution of Ray-tracing Channel Simulation via Attention Mechanism based Deep Learning Model 11 Jan 2023 · 0 repositories · arXiv:2301.04479
-
Dynamic Background Reconstruction via MAE for Infrared Small Target Detection 11 Jan 2023 · 0 repositories · arXiv:2301.04497
-
GPT as Knowledge Worker: A Zero-Shot Evaluation of (AI)CPA Capabilities 11 Jan 2023 · 1 repository · arXiv:2301.04408
-
Head-Free Lightweight Semantic Segmentation with Linear Transformer 11 Jan 2023 · 1 repository · arXiv:2301.04648
-
NarrowBERT: Accelerating Masked Language Model Pretraining and Inference 11 Jan 2023 · 1 repository · arXiv:2301.04761Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Topics in Contextualised Attention Embeddings 11 Jan 2023 · 0 repositories · arXiv:2301.04339
-
Dynamic Grained Encoder for Vision Transformers 10 Jan 2023 · 1 repository · arXiv:2301.03831Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Language Models sounds the Death Knell of Knowledge Graphs 10 Jan 2023 · 0 repositories · arXiv:2301.03980
-
Predicting Hateful Discussions on Reddit using Graph Transformer Networks and Communal Context 10 Jan 2023 · 1 repository · arXiv:2301.04248
-
Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language Models 10 Jan 2023 · 0 repositories · arXiv:2301.03797
-
Streaming Punctuation: A Novel Punctuation Technique Leveraging Bidirectional Context for Continuous Speech Recognition 10 Jan 2023 · 0 repositories · arXiv:2301.03819
-
There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering 10 Jan 2023 · 2 repositories · arXiv:2301.04013
-
Unsupervised Mandarin-Cantonese Machine Translation 10 Jan 2023 · 1 repository · arXiv:2301.03971
-
A Study on the Generality of Neural Network Structures for Monocular Depth Estimation 9 Jan 2023 · 1 repository · arXiv:2301.03169
-
Advances in Medical Image Analysis with Vision Transformers: A Comprehensive Review 9 Jan 2023 · 1 repository · arXiv:2301.03505
-
An Impartial Transformer for Story Visualization 9 Jan 2023 · 0 repositories · arXiv:2301.03563
-
DeMT: Deformable Mixer Transformer for Multi-Task Learning of Dense Prediction 9 Jan 2023 · 1 repository · arXiv:2301.03461Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling 9 Jan 2023 · 2 repositories · arXiv:2301.03580Syntology official (archive's flag): 9 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Logically at Factify 2: A Multi-Modal Fact Checking System Based on Evidence Retrieval techniques and Transformer Encoder Architecture 9 Jan 2023 · 0 repositories · arXiv:2301.03127
-
Online Fake Review Detection Using Supervised Machine Learning And BERT Model 9 Jan 2023 · 0 repositories · arXiv:2301.03225
-
Universal Multimodal Representation for Language Understanding 9 Jan 2023 · 0 repositories · arXiv:2301.03344
-
Automatic Generation of German Drama Texts Using Fine Tuned GPT-2 Models 8 Jan 2023 · 0 repositories · arXiv:2301.03119
-
DeepMatcher: A Deep Transformer-based Network for Robust and Accurate Local Feature Matching 8 Jan 2023 · 1 repository · arXiv:2301.02993
-
HRTransNet: HRFormer-Driven Two-Modality Salient Object Detection 8 Jan 2023 · 1 repository · arXiv:2301.03036
-
App Review Driven Collaborative Bug Finding 7 Jan 2023 · 1 repository · arXiv:2301.02818
-
RLAS-BIABC: A Reinforcement Learning-Based Answer Selection Using the BERT Model Boosted by an Improved ABC Algorithm 7 Jan 2023 · 0 repositories · arXiv:2301.02807
-
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior 6 Jan 2023 · 1 repository · arXiv:2301.02379Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples)
-
Generative Antibody Design for Complementary Chain Pairing Sequences through Encoder-Decoder Language Model 6 Jan 2023 · 0 repositories · arXiv:2301.02748
-
Does compressing activations help model parallel training? 6 Jan 2023 · 0 repositories · arXiv:2301.02654
-
Exploring Efficient Few-shot Adaptation for Vision Transformers 6 Jan 2023 · 1 repository · arXiv:2301.02419
-
Multi-Genre Music Transformer -- Composing Full Length Musical Piece 6 Jan 2023 · 0 repositories · arXiv:2301.02385
-
Systems for Parallel and Distributed Large-Model Deep Learning Training 6 Jan 2023 · 0 repositories · arXiv:2301.02691
-
Adaptive Pattern Extraction Multi-Task Learning for Multi-Step Conversion Estimations 6 Jan 2023 · 0 repositories · arXiv:2301.02494
-
Adaptively Clustering Neighbor Elements for Image-Text Generation 5 Jan 2023 · 1 repository · arXiv:2301.01955
-
CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object Detection 5 Jan 2023 · 0 repositories · arXiv:2301.01970
-
Critical Perspectives: A Benchmark Revealing Pitfalls in PerspectiveAPI 5 Jan 2023 · 1 repository · arXiv:2301.01874
-
Enabling Augmented Segmentation and Registration in Ultrasound-Guided Spinal Surgery via Realistic Ultrasound Synthesis from Diagnostic CT Volume 5 Jan 2023 · 0 repositories · arXiv:2301.01940
-
Learning Feature Recovery Transformer for Occluded Person Re-identification 5 Jan 2023 · 1 repository · arXiv:2301.01879
-
Learning Trajectory-Word Alignments for Video-Language Tasks 5 Jan 2023 · 0 repositories · arXiv:2301.01953
-
Single-round Self-supervised Distributed Learning using Vision Transformer 5 Jan 2023 · 0 repositories · arXiv:2301.02064
-
Scalable Communication for Multi-Agent Reinforcement Learning via Transformer-Based Email Mechanism 5 Jan 2023 · 0 repositories · arXiv:2301.01919
-
Sequentially Controlled Text Generation 5 Jan 2023 · 0 repositories · arXiv:2301.02299
-
Towards Autoformalization of Mathematics and Code Correctness: Experiments with Elementary Proofs 5 Jan 2023 · 1 repository · arXiv:2301.02195Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Towards Long-Term Time-Series Forecasting: Feature, Pattern, and Distribution 5 Jan 2023 · 1 repository · arXiv:2301.02068
-
Extending Source Code Pre-Trained Language Models to Summarise Decompiled Binaries 4 Jan 2023 · 1 repository · arXiv:2301.01701
-
Infomaxformer: Maximum Entropy Transformer for Long Time-Series Forecasting Problem 4 Jan 2023 · 0 repositories · arXiv:2301.01772
-
InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval 4 Jan 2023 · 1 repository · arXiv:2301.01820
-
Multi-Aspect Explainable Inductive Relation Prediction by Sentence Transformer 4 Jan 2023 · 1 repository · arXiv:2301.01664
-
Semi-MAE: Masked Autoencoders for Semi-supervised Vision Transformers 4 Jan 2023 · 0 repositories · arXiv:2301.01431
-
SPTS v2: Single-Point Scene Text Spotting 4 Jan 2023 · 3 repositories · arXiv:2301.01635Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 16 harvested samples) · 3 pointer-only (licence)
-
UniHD at TSAR-2022 Shared Task: Is Compute All We Need for Lexical Simplification? 4 Jan 2023 · 1 repository · arXiv:2301.01764
-
A New Perspective to Boost Vision Transformer for Medical Image Classification 3 Jan 2023 · 0 repositories · arXiv:2301.00989
-
Cross Modal Transformer: Towards Fast and Robust 3D Object Detection 3 Jan 2023 · 2 repositories · arXiv:2301.01283