Methods › General › Attention Modules › Multi-Head Attention › Papers, page 20
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 20 of 249: papers 1,901 to 2,000 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training 15 Feb 2025 · 1 repository · arXiv:2502.10927
-
A synergistic CNN-transformer network with pooling attention fusion for hyperspectral image classification 14 Feb 2025 · 1 repository
-
An Efficient Large Recommendation Model: Towards a Resource-Optimal Scaling Law 14 Feb 2025 · 0 repositories · arXiv:2502.09888
-
An Innovative Next Activity Prediction Approach Using Process Entropy and DAW-Transformer 14 Feb 2025 · 0 repositories · arXiv:2502.10573
-
ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation 14 Feb 2025 · 0 repositories · arXiv:2502.09891
-
Compress image to patches for Vision Transformer 14 Feb 2025 · 1 repository · arXiv:2502.10120
-
Do Large Language Models Reason Causally Like Us? Even Better? 14 Feb 2025 · 0 repositories · arXiv:2502.10215
-
EmbBERT-Q: Breaking Memory Barriers in Embedded NLP 14 Feb 2025 · 0 repositories · arXiv:2502.10001
-
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow 14 Feb 2025 · 0 repositories · arXiv:2502.15765
-
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA 14 Feb 2025 · 0 repositories · arXiv:2502.10497
-
Janus: Collaborative Vision Transformer Under Dynamic Network Environment 14 Feb 2025 · 0 repositories · arXiv:2502.10047
-
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing 14 Feb 2025 · 0 repositories · arXiv:2502.09977
-
Large Language Diffusion Models 14 Feb 2025 · 2 repositories · arXiv:2502.09992Syntology 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Post-training an LLM for RAG? Train on Self-Generated Demonstrations 14 Feb 2025 · 0 repositories · arXiv:2502.10596
-
QMaxViT-Unet+: A Query-Based MaxViT-Unet with Edge Enhancement for Scribble-Supervised Segmentation of Medical Images 14 Feb 2025 · 1 repository · arXiv:2502.10294
-
Simplifying DINO via Coding Rate Regularization 14 Feb 2025 · 0 repositories · arXiv:2502.10385
-
A Hybrid Transformer Model for Fake News Detection: Leveraging Bayesian Optimization and Bidirectional Recurrent Unit 13 Feb 2025 · 0 repositories · arXiv:2502.09097
-
A Physics-Informed Deep Learning Model for MRI Brain Motion Correction 13 Feb 2025 · 1 repository · arXiv:2502.09296
-
Application of Tabular Transformer Architectures for Operating System Fingerprinting 13 Feb 2025 · 1 repository · arXiv:2502.09084
-
Biologically Plausible Brain Graph Transformer 13 Feb 2025 · 1 repository · arXiv:2502.08958Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Can Uniform Meaning Representation Help GPT-4 Translate from Indigenous Languages? 13 Feb 2025 · 0 repositories · arXiv:2502.08900
-
Channel Dependence, Limited Lookback Windows, and the Simplicity of Datasets: How Biased is Time Series Forecasting? 13 Feb 2025 · 0 repositories · arXiv:2502.09683
-
Diverse Transformer Decoding for Offline Reinforcement Learning Using Financial Algorithmic Approaches 13 Feb 2025 · 0 repositories · arXiv:2502.10473
-
E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization 13 Feb 2025 · 0 repositories · arXiv:2502.09164
-
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables 13 Feb 2025 · 0 repositories · arXiv:2502.09073
-
Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification 13 Feb 2025 · 0 repositories · arXiv:2502.08997
-
Improving TCM Question Answering through Tree-Organized Self-Reflective Retrieval with LLMs 13 Feb 2025 · 0 repositories · arXiv:2502.09156
-
KIMAs: A Configurable Knowledge Integrated Multi-Agent System 13 Feb 2025 · 0 repositories · arXiv:2502.09596
-
MC2SleepNet: Multi-modal Cross-masking with Contrastive Learning for Sleep Stage Classification 13 Feb 2025 · 1 repository · arXiv:2502.17470
-
Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning 13 Feb 2025 · 0 repositories · arXiv:2502.09022
-
Predicting Cognitive Decline: A Multimodal AI Approach to Dementia Screening from Speech 13 Feb 2025 · 0 repositories · arXiv:2502.08862
-
Residual Transformer Fusion Network for Salt and Pepper Image Denoising 13 Feb 2025 · 0 repositories · arXiv:2502.09000
-
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? 13 Feb 2025 · 1 repository · arXiv:2502.09416Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues 13 Feb 2025 · 0 repositories · arXiv:2502.09120
-
Utilizing Pre-trained and Large Language Models for 10-K Items Segmentation 13 Feb 2025 · 0 repositories · arXiv:2502.08875
-
What exactly has TabPFN learned to do? 13 Feb 2025 · 1 repository · arXiv:2502.08978
-
A Survey on Image Quality Assessment: Insights, Analysis, and Future Outlook 12 Feb 2025 · 0 repositories · arXiv:2502.08540
-
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation 12 Feb 2025 · 1 repository · arXiv:2502.08826
-
TANTE: Time-Adaptive Operator Learning via Neural Taylor Expansion 12 Feb 2025 · 0 repositories · arXiv:2502.08574
-
Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning 12 Feb 2025 · 0 repositories · arXiv:2502.08482
-
Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance 12 Feb 2025 · 1 repository · arXiv:2502.08127Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
HDT: Hierarchical Discrete Transformer for Multivariate Time Series Forecasting 12 Feb 2025 · 1 repository · arXiv:2502.08302
-
Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation 12 Feb 2025 · 1 repository · arXiv:2502.08347
-
InTAR: Inter-Task Auto-Reconfigurable Accelerator Design for High Data Volume Variation in DNNs 12 Feb 2025 · 1 repository · arXiv:2502.08807
-
ParetoRAG: Leveraging Sentence-Context Attention for Robust and Efficient Retrieval-Augmented Generation 12 Feb 2025 · 0 repositories · arXiv:2502.08178
-
Rethinking Tokenized Graph Transformers for Node Classification 12 Feb 2025 · 0 repositories · arXiv:2502.08101
-
Self-Evaluation for Job-Shop Scheduling 12 Feb 2025 · 0 repositories · arXiv:2502.08684
-
Systematic Knowledge Injection into Large Language Models via Diverse Augmentation for Domain-Specific RAG 12 Feb 2025 · 1 repository · arXiv:2502.08356Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
YNote: A Novel Music Notation for Fine-Tuning LLMs in Music Generation 12 Feb 2025 · 0 repositories · arXiv:2502.10467
-
5D Neural Surrogates for Nonlinear Gyrokinetic Simulations of Plasma Turbulence 11 Feb 2025 · 0 repositories · arXiv:2502.07469
-
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases 11 Feb 2025 · 1 repository · arXiv:2502.07188
-
An Advanced NLP Framework for Automated Medical Diagnosis with DeBERTa and Dynamic Contextual Positional Gating 11 Feb 2025 · 0 repositories · arXiv:2502.07755
-
Auditing Prompt Caching in Language Model APIs 11 Feb 2025 · 1 repository · arXiv:2502.07776Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Automated Capability Discovery via Model Self-Exploration 11 Feb 2025 · 2 repositories · arXiv:2502.07577
-
CausalGeD: Blending Causality and Diffusion for Spatial Gene Expression Generation 11 Feb 2025 · 0 repositories · arXiv:2502.07751
-
Dataset Ownership Verification in Contrastive Pre-trained Models 11 Feb 2025 · 1 repository · arXiv:2502.07276
-
Deep Semantic Graph Learning via LLM based Node Enhancement 11 Feb 2025 · 0 repositories · arXiv:2502.07982
-
Dense Object Detection Based on De-homogenized Queries 11 Feb 2025 · 0 repositories · arXiv:2502.07194
-
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving 11 Feb 2025 · 0 repositories · arXiv:2502.07417
-
FoQA: A Faroese Question-Answering Dataset 11 Feb 2025 · 0 repositories · arXiv:2502.07642
-
Grammar Control in Dialogue Response Generation for Language Learning Chatbots 11 Feb 2025 · 1 repository · arXiv:2502.07544
-
Graph RAG-Tool Fusion 11 Feb 2025 · 1 repository · arXiv:2502.07223
-
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting 11 Feb 2025 · 1 repository · arXiv:2502.07244Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
MAAT: Mamba Adaptive Anomaly Transformer with association discrepancy for time series 11 Feb 2025 · 1 repository · arXiv:2502.07858
-
Making Language Models Robust Against Negation 11 Feb 2025 · 0 repositories · arXiv:2502.07717
-
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More 11 Feb 2025 · 1 repository · arXiv:2502.07490
-
MIGT: Memory Instance Gated Transformer Framework for Financial Portfolio Management 11 Feb 2025 · 0 repositories · arXiv:2502.07280
-
OpenGrok: Enhancing SNS Data Processing with Distilled Knowledge and Mask-like Mechanisms 11 Feb 2025 · 1 repository · arXiv:2502.07312
-
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers 11 Feb 2025 · 0 repositories · arXiv:2502.07436
-
Tractable Transformers for Flexible Conditional Generation 11 Feb 2025 · 0 repositories · arXiv:2502.07616
-
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation 11 Feb 2025 · 0 repositories · arXiv:2502.07531
-
A Simple yet Effective DDG Predictor is An Unsupervised Antibody Optimizer and Explainer 10 Feb 2025 · 1 repository · arXiv:2502.06913
-
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation 10 Feb 2025 · 0 repositories · arXiv:2502.06205Syntology 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
ConMeC: A Dataset for Metonymy Resolution with Common Nouns 10 Feb 2025 · 1 repository · arXiv:2502.06087
-
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models 10 Feb 2025 · 0 repositories · arXiv:2502.06279
-
Find Central Dogma Again: Leveraging Multilingual Transfer in Large Language Models 10 Feb 2025 · 0 repositories · arXiv:2502.06253
-
Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI 10 Feb 2025 · 0 repositories · arXiv:2502.07017
-
Foundation Model of Electronic Medical Records for Adaptive Risk Estimation 10 Feb 2025 · 1 repository · arXiv:2502.06124
-
Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing 10 Feb 2025 · 0 repositories · arXiv:2502.06219
-
History-Guided Video Diffusion 10 Feb 2025 · 1 repository · arXiv:2502.06764
-
LLMs in Software Security: A Survey of Vulnerability Detection Techniques and Insights 10 Feb 2025 · 0 repositories · arXiv:2502.07049
-
Leveraging GPT-4o Efficiency for Detecting Rework Anomaly in Business Processes 10 Feb 2025 · 0 repositories · arXiv:2502.06918
-
Multimodal Task Representation Memory Bank vs. Catastrophic Forgetting in Anomaly Detection 10 Feb 2025 · 0 repositories · arXiv:2502.06194
-
Optimizing Knowledge Integration in Retrieval-Augmented Generation with Self-Selection 10 Feb 2025 · 0 repositories · arXiv:2502.06148
-
Powerformer: A Transformer with Weighted Causal Attention for Time-series Forecasting 10 Feb 2025 · 1 repository · arXiv:2502.06151
-
RALLRec: Improving Retrieval Augmented Large Language Model Recommendation with Representation Learning 10 Feb 2025 · 1 repository · arXiv:2502.06101
-
Towards bandit-based prompt-tuning for in-the-wild foundation agents 10 Feb 2025 · 0 repositories · arXiv:2502.06358
-
Towards Copyright Protection for Knowledge Bases of Retrieval-augmented Language Models via Reasoning 10 Feb 2025 · 0 repositories · arXiv:2502.10440
-
Unconstrained Body Recognition at Altitude and Range: Comparing Four Approaches 10 Feb 2025 · 0 repositories · arXiv:2502.07130
-
Utilizing Novelty-based Evolution Strategies to Train Transformers in Reinforcement Learning 10 Feb 2025 · 0 repositories · arXiv:2502.06301
-
ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models 10 Feb 2025 · 0 repositories · arXiv:2502.06741
-
Benchmarking Prompt Engineering Techniques for Secure Code Generation with GPT Models 9 Feb 2025 · 0 repositories · arXiv:2502.06039
-
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training 9 Feb 2025 · 0 repositories · arXiv:2502.06902
-
Enhancing Financial Time-Series Forecasting with Retrieval-Augmented Large Language Models 9 Feb 2025 · 0 repositories · arXiv:2502.05878
-
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition 9 Feb 2025 · 0 repositories
-
Investigating Compositional Reasoning in Time Series Foundation Models 9 Feb 2025 · 1 repository · arXiv:2502.06037
-
Large Language Models for In-File Vulnerability Localization Can Be "Lost in the End" 9 Feb 2025 · 0 repositories · arXiv:2502.06898
-
LM2: Large Memory Models 9 Feb 2025 · 2 repositories · arXiv:2502.06049
-
Saving 77% of the Parameters in Large Language Models Technical Report 9 Feb 2025 · 1 repository
-
ScaffoldGPT: A Scaffold-based GPT Model for Drug Optimization 9 Feb 2025 · 0 repositories · arXiv:2502.06891