Methods › General › Attention Modules › Multi-Head Attention › Papers, page 14
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 14 of 249: papers 1,301 to 1,400 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Shushing! Let's Imagine an Authentic Speech from the Silent Video 19 Mar 2025 · 0 repositories · arXiv:2503.14928
-
TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification 19 Mar 2025 · 0 repositories · arXiv:2503.15289
-
TruthLens:A Training-Free Paradigm for DeepFake Detection 19 Mar 2025 · 0 repositories · arXiv:2503.15342
-
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study 19 Mar 2025 · 1 repository · arXiv:2503.15579
-
A-SCoRe: Attention-based Scene Coordinate Regression for wide-ranging scenarios 18 Mar 2025 · 1 repository · arXiv:2503.13982
-
Binary AddiVortes: (Bayesian) Additive Voronoi Tessellations for Binary Classification with an application to Predicting Home Mortgage Application Outcomes 18 Mar 2025 · 0 repositories · arXiv:2503.21792
-
BurTorch: Revisiting Training from First Principles by Coupling Autodiff, Math Optimization, and Systems 18 Mar 2025 · 1 repository · arXiv:2503.13795
-
CTSAC: Curriculum-Based Transformer Soft Actor-Critic for Goal-Oriented Robot Exploration 18 Mar 2025 · 0 repositories · arXiv:2503.14254
-
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer 18 Mar 2025 · 1 repository · arXiv:2503.14640
-
Enhancing LLM Generation with Knowledge Hypergraph for Evidence-Based Medicine 18 Mar 2025 · 0 repositories · arXiv:2503.16530
-
Fast Autoregressive Video Generation with Diagonal Decoding 18 Mar 2025 · 0 repositories · arXiv:2503.14070
-
Good/Evil Reputation Judgment of Celebrities by LLMs via Retrieval Augmented Generation 18 Mar 2025 · 0 repositories · arXiv:2503.14382
-
Gricean Norms as a Basis for Effective Collaboration 18 Mar 2025 · 1 repository · arXiv:2503.14484
-
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System 18 Mar 2025 · 1 repository · arXiv:2503.14258
-
Beyond Single Pass, Looping Through Time: KG-IRAG with Iterative Knowledge Retrieval 18 Mar 2025 · 0 repositories · arXiv:2503.14234
-
Large Language Models for Virtual Human Gesture Selection 18 Mar 2025 · 0 repositories · arXiv:2503.14408
-
MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding 18 Mar 2025 · 1 repository · arXiv:2503.13964Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
MoK-RAG: Mixture of Knowledge Paths Enhanced Retrieval-Augmented Generation for Embodied AI Environments 18 Mar 2025 · 0 repositories · arXiv:2503.13882
-
Multimodal Feature-Driven Deep Learning for the Prediction of Duck Body Dimensions and Weight 18 Mar 2025 · 0 repositories · arXiv:2503.14001
-
PENCIL: Long Thoughts with Short Memory 18 Mar 2025 · 1 repository · arXiv:2503.14337Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Predicting Human Choice Between Textually Described Lotteries 18 Mar 2025 · 0 repositories · arXiv:2503.14004
-
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving 18 Mar 2025 · 1 repository · arXiv:2503.14649
-
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking 18 Mar 2025 · 0 repositories · arXiv:2503.13805
-
Theoretical Foundation of Flow-Based Time Series Generation: Provable Approximation, Generalization, and Efficiency 18 Mar 2025 · 0 repositories · arXiv:2503.14076
-
XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants 18 Mar 2025 · 0 repositories · arXiv:2503.14281
-
A Reinforcement Learning-Driven Transformer GAN for Molecular Generation 17 Mar 2025 · 0 repositories · arXiv:2503.12796
-
A Survey on Transformer Context Extension: Approaches and Evaluation 17 Mar 2025 · 0 repositories · arXiv:2503.13299
-
Advancing Chronic Tuberculosis Diagnostics Using Vision-Language Models: A Multi modal Framework for Precision Analysis 17 Mar 2025 · 0 repositories · arXiv:2503.14536
-
An interpretable approach to automating the assessment of biofouling in video footage 17 Mar 2025 · 1 repository · arXiv:2503.12875
-
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs 17 Mar 2025 · 0 repositories · arXiv:2503.13149
-
Can Language Models Follow Multiple Turns of Entangled Instructions? 17 Mar 2025 · 1 repository · arXiv:2503.13222
-
Feature Extraction and Analysis for GPT-Generated Text 17 Mar 2025 · 0 repositories · arXiv:2503.13687
-
Generative AI for Software Architecture. Applications, Trends, Challenges, and Future Directions 17 Mar 2025 · 0 repositories · arXiv:2503.13310
-
Humanoid Policy ~ Human Policy 17 Mar 2025 · 0 repositories · arXiv:2503.13441
-
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention 17 Mar 2025 · 1 repository · arXiv:2503.12734
-
MES-RAG: Bringing Multi-modal, Entity-Storage, and Secure Enhancements to RAG 17 Mar 2025 · 1 repository · arXiv:2503.13563
-
OSCAR: Online Soft Compression And Reranking 17 Mar 2025 · 0 repositories · arXiv:2504.07109
-
PAUSE: Low-Latency and Privacy-Aware Active User Selection for Federated Learning 17 Mar 2025 · 1 repository · arXiv:2503.13173
-
Privacy-Aware RAG: Secure and Isolated Knowledge Retrieval 17 Mar 2025 · 0 repositories · arXiv:2503.15548
-
SeisRDT: Latent Diffusion Model Based On Representation Learning For Seismic Data Interpolation And Reconstruction 17 Mar 2025 · 0 repositories · arXiv:2503.21791
-
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data 17 Mar 2025 · 0 repositories · arXiv:2503.12843
-
Fourier-Based 3D Multistage Transformer for Aberration Correction in Multicellular Specimens 16 Mar 2025 · 2 repositories · arXiv:2503.12593
-
Fragile Mastery: Are Domain-Specific Trade-Offs Undermining On-Device Language Models? 16 Mar 2025 · 0 repositories · arXiv:2503.22698
-
GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation 16 Mar 2025 · 0 repositories · arXiv:2503.12600
-
Semantic Matters: Multimodal Features for Affective Analysis 16 Mar 2025 · 0 repositories · arXiv:2504.11460
-
PA-CFL: Privacy-Adaptive Clustered Federated Learning for Transformer-Based Sales Forecasting on Heterogeneous Retail Data 15 Mar 2025 · 0 repositories · arXiv:2503.12220
-
Changing Base Without Losing Pace: A GPU-Efficient Alternative to MatMul in DNNs 15 Mar 2025 · 0 repositories · arXiv:2503.12211
-
Fast Critical Clearing Time Calculation for Power Systems with Synchronous and Asynchronous Generation 15 Mar 2025 · 0 repositories · arXiv:2503.12132
-
Integrating Chain-of-Thought and Retrieval Augmented Generation Enhances Rare Disease Diagnosis from Clinical Notes 15 Mar 2025 · 0 repositories · arXiv:2503.12286
-
Language Models for Automated Classification of Brain MRI Reports and Growth Chart Generation 15 Mar 2025 · 0 repositories · arXiv:2503.12143
-
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks 15 Mar 2025 · 1 repository · arXiv:2504.03665
-
Maritime Mission Planning for Unmanned Surface Vessel using Large Language Model 15 Mar 2025 · 0 repositories · arXiv:2503.12065
-
Addressing Information Loss and Interaction Collapse: A Dual Enhanced Attention Framework for Feature Interaction 14 Mar 2025 · 0 repositories · arXiv:2503.11233
-
Alzheimer's Disease Classification Using Retinal OCT: TransnetOCT and Swin Transformer Models 14 Mar 2025 · 0 repositories · arXiv:2503.11511
-
Asynchronous Sharpness-Aware Minimization For Fast and Accurate Deep Learning 14 Mar 2025 · 0 repositories · arXiv:2503.11147
-
Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation 14 Mar 2025 · 0 repositories · arXiv:2503.11096
-
BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model 14 Mar 2025 · 1 repository · arXiv:2503.11372
-
Combining Causal Models for More Accurate Abstractions of Neural Networks 14 Mar 2025 · 1 repository · arXiv:2503.11429Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Context-Aware Rule Mining Using a Dynamic Transformer-Based Framework 14 Mar 2025 · 0 repositories · arXiv:2503.11125
-
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models 14 Mar 2025 · 0 repositories · arXiv:2503.11265
-
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment 14 Mar 2025 · 0 repositories · arXiv:2503.11229
-
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding 14 Mar 2025 · 0 repositories · arXiv:2503.11108
-
MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery 14 Mar 2025 · 0 repositories · arXiv:2503.11219
-
Prompt Sentiment: The Catalyst for LLM Change 14 Mar 2025 · 0 repositories · arXiv:2503.13510
-
RAG-KG-IL: A Multi-Agent Hybrid Framework for Reducing Hallucinations and Enhancing LLM Reasoning through RAG and Incremental Knowledge Graph Learning Integration 14 Mar 2025 · 0 repositories · arXiv:2503.13514
-
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking 14 Mar 2025 · 2 repositories · arXiv:2504.07104
-
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation 14 Mar 2025 · 0 repositories · arXiv:2503.11348
-
Semantic and Contextual Modeling for Malicious Comment Detection with BERT-BiLSTM 14 Mar 2025 · 0 repositories · arXiv:2503.11084
-
Solution for 8th Competition on Affective & Behavior Analysis in-the-wild 14 Mar 2025 · 0 repositories · arXiv:2503.11115
-
Text Compression for Efficient Language Generation 14 Mar 2025 · 0 repositories · arXiv:2503.11426
-
TransiT: Transient Transformer for Non-line-of-sight Videography 14 Mar 2025 · 0 repositories · arXiv:2503.11328
-
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing 14 Mar 2025 · 1 repository · arXiv:2503.11629Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 6 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective 14 Mar 2025 · 1 repository · arXiv:2503.11272
-
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 13 Mar 2025 · 1 repository · arXiv:2503.10635Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization 13 Mar 2025 · 0 repositories · arXiv:2503.10354
-
Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM 13 Mar 2025 · 0 repositories · arXiv:2503.10071
-
ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents 13 Mar 2025 · 0 repositories · arXiv:2503.10233
-
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation 13 Mar 2025 · 0 repositories · arXiv:2503.10720
-
AudioX: Diffusion Transformer for Anything-to-Audio Generation 13 Mar 2025 · 0 repositories · arXiv:2503.10522
-
ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models 13 Mar 2025 · 0 repositories · arXiv:2503.10937
-
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception 13 Mar 2025 · 0 repositories · arXiv:2503.13504
-
Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text 13 Mar 2025 · 1 repository · arXiv:2503.10095
-
Compositional Subspace Representation Fine-tuning for Adaptive Large Language Models 13 Mar 2025 · 0 repositories · arXiv:2503.10617
-
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers 13 Mar 2025 · 0 repositories · arXiv:2503.09942
-
CountPath: Automating Fragment Counting in Digital Pathology 13 Mar 2025 · 0 repositories · arXiv:2503.10520
-
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark 13 Mar 2025 · 0 repositories · arXiv:2503.10357
-
Edge-Fog Computing-Enabled EEG Data Compression via Asymmetrical Variational Discrete Cosine Transform Network 13 Mar 2025 · 0 repositories · arXiv:2503.09961
-
Emotion Recognition with CLIP and Sequential Learning 13 Mar 2025 · 0 repositories · arXiv:2503.09929
-
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG 13 Mar 2025 · 1 repository · arXiv:2504.07103
-
Fixed-Point RNNs: From Diagonal to Dense in a Few Iterations 13 Mar 2025 · 0 repositories · arXiv:2503.10799
-
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding 13 Mar 2025 · 0 repositories · arXiv:2503.10135Syntology 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 8 pointer-only (licence)
-
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education 13 Mar 2025 · 0 repositories · arXiv:2503.13508
-
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs 13 Mar 2025 · 0 repositories · arXiv:2503.10337
-
Multi-Domain Biometric Recognition using Body Embeddings 13 Mar 2025 · 0 repositories · arXiv:2503.10931
-
Predicting Stock Movement with BERTweet and Transformers 13 Mar 2025 · 0 repositories · arXiv:2503.10957
-
Radar: Fast Long-Context Decoding for Any Transformer 13 Mar 2025 · 1 repository · arXiv:2503.10571Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Retrieval-Augmented Generation with Hierarchical Knowledge 13 Mar 2025 · 1 repository · arXiv:2503.10150Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Robustness Tokens: Towards Adversarial Robustness of Transformers 13 Mar 2025 · 1 repository · arXiv:2503.10191
-
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search 13 Mar 2025 · 0 repositories · arXiv:2503.10619
-
TacticExpert: Spatial-Temporal Graph Language Model for Basketball Tactics 13 Mar 2025 · 0 repositories · arXiv:2503.10722