Methods › General › Attention Modules › Multi-Head Attention › Papers, page 115
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 115 of 249: papers 11,401 to 11,500 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Does the "most sinfully decadent cake ever" taste good? Answering Yes/No Questions from Figurative Contexts 24 Sep 2023 · 0 repositories · arXiv:2309.13748
-
Global-correlated 3D-decoupling Transformer for Clothed Avatar Reconstruction 24 Sep 2023 · 1 repository · arXiv:2309.13524Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
MediViSTA: Medical Video Segmentation via Temporal Fusion SAM Adaptation for Echocardiography 24 Sep 2023 · 1 repository · arXiv:2309.13539
-
MOSAIC: Multi-Object Segmented Arbitrary Stylization Using CLIP 24 Sep 2023 · 0 repositories · arXiv:2309.13716
-
Multi-Dimensional Hyena for Spatial Inductive Bias 24 Sep 2023 · 0 repositories · arXiv:2309.13600
-
Natural Language based Context Modeling and Reasoning for Ubiquitous Computing with Large Language Models: A Tutorial 24 Sep 2023 · 0 repositories · arXiv:2309.15074
-
Seeing Is Not Always Believing: Invisible Collision Attack and Defence on Pre-Trained Models 24 Sep 2023 · 1 repository · arXiv:2309.13579
-
A Chat About Boring Problems: Studying GPT-based text normalization 23 Sep 2023 · 0 repositories · arXiv:2309.13426
-
Algorithms for Object Detection in Substations 23 Sep 2023 · 0 repositories · arXiv:2311.07577
-
Asca: less audio data is more insightful 23 Sep 2023 · 1 repository · arXiv:2309.13373
-
Attention Is All You Need For Blind Room Volume Estimation 23 Sep 2023 · 0 repositories · arXiv:2309.13504
-
EMGTFNet: Fuzzy Vision Transformer to decode Upperlimb sEMG signals for Hand Gestures Recognition 23 Sep 2023 · 0 repositories · arXiv:2310.03754
-
Probing the Moral Development of Large Language Models through Defining Issues Test 23 Sep 2023 · 0 repositories · arXiv:2309.13356
-
GlotScript: A Resource and Tool for Low Resource Writing System Identification 23 Sep 2023 · 1 repository · arXiv:2309.13320Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Hindi to English: Transformer-Based Neural Machine Translation 23 Sep 2023 · 0 repositories · arXiv:2309.13222
-
Lexical Squad@Multimodal Hate Speech Event Detection 2023: Multimodal Hate Speech Detection using Fused Ensemble Approach 23 Sep 2023 · 1 repository · arXiv:2309.13354
-
RBFormer: Improve Adversarial Robustness of Transformer by Robust Bias 23 Sep 2023 · 0 repositories · arXiv:2309.13245
-
UniHead: Unifying Multi-Perception for Detection Heads 23 Sep 2023 · 1 repository · arXiv:2309.13242
-
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer 22 Sep 2023 · 1 repository · arXiv:2309.12689
-
AntiBARTy Diffusion for Property Guided Antibody Design 22 Sep 2023 · 0 repositories · arXiv:2309.13129
-
Associative Transformer 22 Sep 2023 · 1 repository · arXiv:2309.12862
-
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP 22 Sep 2023 · 0 repositories · arXiv:2309.13173
-
ClusterFormer: Clustering As A Universal Visual Learner 22 Sep 2023 · 1 repository · arXiv:2309.13196
-
Contextual Emotion Estimation from Image Captions 22 Sep 2023 · 0 repositories · arXiv:2309.13136
-
DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single Image 22 Sep 2023 · 0 repositories · arXiv:2309.12594
-
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis 22 Sep 2023 · 0 repositories · arXiv:2309.12792
-
Investigating Large Language Models and Control Mechanisms to Improve Text Readability of Biomedical Abstracts 22 Sep 2023 · 1 repository · arXiv:2309.13202
-
Large Language Models Are Also Good Prototypical Commonsense Reasoners 22 Sep 2023 · 0 repositories · arXiv:2309.13165
-
Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where 22 Sep 2023 · 1 repository · arXiv:2309.12757
-
Modeling Spatiotemporal Periodicity and Collaborative Signal for Local-Life Service Recommendation 22 Sep 2023 · 0 repositories · arXiv:2309.12565
-
PointSSC: A Cooperative Vehicle-Infrastructure Point Cloud Benchmark for Semantic Scene Completion 22 Sep 2023 · 1 repository · arXiv:2309.12708Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs 22 Sep 2023 · 2 repositories · arXiv:2309.13007Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SPION: Layer-Wise Sparse Training of Transformer via Convolutional Flood Filling 22 Sep 2023 · 0 repositories · arXiv:2309.12578
-
StyloMetrix: An Open-Source Multilingual Tool for Representing Stylometric Vectors 22 Sep 2023 · 1 repository · arXiv:2309.12810
-
TOPFORMER: Topology-Aware Authorship Attribution of Deepfake Texts with Diverse Writing Styles 22 Sep 2023 · 1 repository · arXiv:2309.12934
-
TrTr: A Versatile Pre-Trained Large Traffic Model based on Transformer for Capturing Trajectory Diversity in Vehicle Population 22 Sep 2023 · 0 repositories · arXiv:2309.12677
-
Vision Transformers for Computer Go 22 Sep 2023 · 0 repositories · arXiv:2309.12675
-
Goal-Oriented Prompt Attack and Safety Evaluation for LLMs 21 Sep 2023 · 2 repositories · arXiv:2309.11830
-
A Robust and Opponent-Aware League Training Method for StarCraft II 21 Sep 2023 · 0 repositories
-
AceGPT, Localizing Large Language Models in Arabic 21 Sep 2023 · 1 repository · arXiv:2309.12053
-
Adaptive Input-image Normalization for Solving the Mode Collapse Problem in GAN-based X-ray Images 21 Sep 2023 · 0 repositories · arXiv:2309.12245
-
Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection 21 Sep 2023 · 1 repository · arXiv:2309.12247
-
BadTrack: A Poison-Only Backdoor Attack on Visual Object Tracking 21 Sep 2023 · 0 repositories
-
BayesTune: Bayesian Sparse Deep Model Fine-tuning 21 Sep 2023 · 1 repository
-
BIOT: Biosignal Transformer for Cross-data Learning in the Wild 21 Sep 2023 · 1 repository
-
Blockwise Parallel Transformers for Large Context Models 21 Sep 2023 · 1 repository
-
Boolformer: Symbolic Regression of Logic Functions with Transformers 21 Sep 2023 · 1 repository · arXiv:2309.12207
-
Can LLMs Augment Low-Resource Reading Comprehension Datasets? Opportunities and Challenges 21 Sep 2023 · 0 repositories · arXiv:2309.12426
-
Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer Models 21 Sep 2023 · 1 repository
-
ClusterFomer: Clustering As A Universal Visual Learner 21 Sep 2023 · 1 repository
-
Code Soliloquies for Accurate Calculations in Large Language Models 21 Sep 2023 · 1 repository · arXiv:2309.12161Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Conformal PID Control for Time Series Prediction 21 Sep 2023 · 1 repository
-
Constraints First: A New MDD-based Model to Generate Sentences Under Constraints 21 Sep 2023 · 0 repositories · arXiv:2309.12415
-
Convolution and Attention Mixer for Synthetic Aperture Radar Image Change Detection 21 Sep 2023 · 1 repository · arXiv:2309.12010Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
DAC-DETR: Divide the Attention Layers and Conquer 21 Sep 2023 · 1 repository
-
DEYOv3: DETR with YOLO for Real-time Object Detection 21 Sep 2023 · 0 repositories · arXiv:2309.11851
-
DualToken-ViT: Position-aware Efficient Vision Transformer with Dual Token Fusion 21 Sep 2023 · 0 repositories · arXiv:2309.12424
-
FLSL: Feature-level Self-supervised Learning 21 Sep 2023 · 1 repository
-
Fully Transformer-Equipped Architecture for End-to-End Referring Video Object Segmentation 21 Sep 2023 · 0 repositories · arXiv:2309.11933
-
Geometric Transformer with Interatomic Positional Encoding 21 Sep 2023 · 1 repository
-
H3T: Efficient Integration of Memory Optimization and Parallelism for Large-scale Transformer Training 21 Sep 2023 · 1 repository
-
Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 21 Sep 2023 · 1 repository
-
LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer 21 Sep 2023 · 1 repository
-
LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction 21 Sep 2023 · 0 repositories
-
LLMR: Real-time Prompting of Interactive Worlds using Large Language Models 21 Sep 2023 · 0 repositories · arXiv:2309.12276
-
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset 21 Sep 2023 · 5 repositories · arXiv:2309.11998
-
Making Scalable Meta Learning Practical 21 Sep 2023 · 1 repository
-
Marich: A Query-efficient Distributionally Equivalent Model Extraction Attack 21 Sep 2023 · 1 repository
-
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models 21 Sep 2023 · 1 repository · arXiv:2309.12284Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 14 where Syntology's instrument failed) · 7 unverified (of 22 harvested samples)
-
MG-ViT: A Multi-Granularity Method for Compact and Efficient Vision Transformers 21 Sep 2023 · 0 repositories
-
MiChao-HuaFen 1.0: A Specialized Pre-trained Corpus Dataset for Domain-specific Large Models 21 Sep 2023 · 0 repositories · arXiv:2309.13079
-
Mitigating Over-smoothing in Transformers via Regularized Nonlocal Functionals 21 Sep 2023 · 0 repositories
-
Multimodal Deep Learning for Scientific Imaging Interpretation 21 Sep 2023 · 0 repositories · arXiv:2309.12460
-
Neural Data Transformer 2: Multi-context Pretraining for Neural Spiking Activity 21 Sep 2023 · 1 repository
-
On the Relationship between Skill Neurons and Robustness in Prompt Tuning 21 Sep 2023 · 1 repository · arXiv:2309.12263
-
One Fits All: Power General Time Series Analysis by Pretrained LM 21 Sep 2023 · 2 repositories
-
OSNet & MNetO: Two Types of General Reconstruction Architectures for Linear Computed Tomography in Multi-Scenarios 21 Sep 2023 · 0 repositories · arXiv:2309.11858
-
PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation 21 Sep 2023 · 1 repository · arXiv:2309.12303
-
Patch n’ Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution 21 Sep 2023 · 0 repositories
-
QuantSR: Accurate Low-bit Quantization for Efficient Image Super-Resolution 21 Sep 2023 · 1 repository
-
Random-Access Infinite Context Length for Transformers 21 Sep 2023 · 1 repository
-
[Re] Exploring the Role of Grammar and Word Choice in Bias Toward African American English (AAE) in Hate Speech Classification 21 Sep 2023 · 0 repositories
-
[Re] Masked Autoencoders Are Small Scale Vision Learners: A Reproduction Under Resource Constraints 21 Sep 2023 · 1 repository
-
[Re] On the Reproducibility of CartoonX 21 Sep 2023 · 0 repositories
-
REFINE: A Fine-Grained Medication Recommendation System Using Deep Learning and Personalized Drug Interaction Modeling 21 Sep 2023 · 0 repositories
-
SLHCat: Mapping Wikipedia Categories and Lists to DBpedia by Leveraging Semantic, Lexical, and Hierarchical Features 21 Sep 2023 · 0 repositories · arXiv:2309.11791
-
Spatial-Temporal Transformer based Video Compression Framework 21 Sep 2023 · 0 repositories · arXiv:2309.11913
-
SPICED: News Similarity Detection Dataset with Multiple Topics and Complexity Levels 21 Sep 2023 · 0 repositories · arXiv:2309.13080
-
SPRING: Studying Papers and Reasoning to play Games 21 Sep 2023 · 0 repositories
-
Stock Market Sentiment Classification and Backtesting via Fine-tuned BERT 21 Sep 2023 · 0 repositories · arXiv:2309.11979
-
TART: A plug-and-play Transformer module for task-agnostic reasoning 21 Sep 2023 · 1 repository
-
The Cambridge Law Corpus: A Dataset for Legal AI Research 21 Sep 2023 · 0 repositories · arXiv:2309.12269
-
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" 21 Sep 2023 · 2 repositories · arXiv:2309.12288Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
TOA: Task-oriented Active VQA 21 Sep 2023 · 0 repositories
-
Toward Re-Identifying Any Animal 21 Sep 2023 · 0 repositories
-
Towards Efficient Pre-Trained Language Model via Feature Correlation Distillation 21 Sep 2023 · 0 repositories
-
Unified 3D Segmenter As Prototypical Classifiers 21 Sep 2023 · 1 repository
-
A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models 20 Sep 2023 · 1 repository · arXiv:2309.11674Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
AttentionMix: Data augmentation method that relies on BERT attention mechanism 20 Sep 2023 · 0 repositories · arXiv:2309.11104
-
Automatic Bat Call Classification using Transformer Networks 20 Sep 2023 · 0 repositories · arXiv:2309.11218