Methods › General › Attention Modules › Multi-Head Attention › Papers, page 4
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 4 of 249: papers 301 to 400 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty 22 May 2025 · 0 repositories · arXiv:2505.17281
-
Swin Transformer for Robust CGI Images Detection: Intra- and Inter-Dataset Analysis across Multiple Color Spaces 22 May 2025 · 0 repositories · arXiv:2505.16253
-
The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm 22 May 2025 · 0 repositories · arXiv:2505.16932
-
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers 22 May 2025 · 0 repositories · arXiv:2505.16241
-
Training-Free Efficient Video Generation via Dynamic Token Carving 22 May 2025 · 1 repository · arXiv:2505.16864
-
Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning 22 May 2025 · 1 repository · arXiv:2505.16270Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Understanding Differential Transformer Unchains Pretrained Self-Attentions 22 May 2025 · 0 repositories · arXiv:2505.16333
-
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering 22 May 2025 · 0 repositories · arXiv:2505.17326
-
Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks 22 May 2025 · 1 repository · arXiv:2505.16849
-
AdUE: Improving uncertainty estimation head for LoRA adapters in LLMs 21 May 2025 · 0 repositories · arXiv:2505.15443
-
An Efficient Private GPT Never Autoregressively Decodes 21 May 2025 · 0 repositories · arXiv:2505.15252
-
An Exploratory Approach Towards Investigating and Explaining Vision Transformer and Transfer Learning for Brain Disease Detection 21 May 2025 · 0 repositories · arXiv:2505.16039
-
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems 21 May 2025 · 0 repositories · arXiv:2505.15216
-
BR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law 21 May 2025 · 0 repositories · arXiv:2505.15916
-
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 21 May 2025 · 0 repositories · arXiv:2505.15045
-
Domain Adaptive Skin Lesion Classification via Conformal Ensemble of Vision Transformers 21 May 2025 · 0 repositories · arXiv:2505.15997
-
Filtering Learning Histories Enhances In-Context Reinforcement Learning 21 May 2025 · 0 repositories · arXiv:2505.15143
-
HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases 21 May 2025 · 1 repository · arXiv:2505.15701
-
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation 21 May 2025 · 0 repositories · arXiv:2505.15872
-
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition 21 May 2025 · 0 repositories · arXiv:2505.15192
-
Leveraging Large Language Models for Command Injection Vulnerability Analysis in Python: An Empirical Study on Popular Open-Source Projects 21 May 2025 · 0 repositories · arXiv:2505.15088
-
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming 21 May 2025 · 0 repositories · arXiv:2505.15039Syntology 11 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 8 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation 21 May 2025 · 0 repositories · arXiv:2505.15696
-
Mechanistic Insights into Grokking from the Embedding Layer 21 May 2025 · 0 repositories · arXiv:2505.15624
-
Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains 21 May 2025 · 0 repositories · arXiv:2505.16014
-
Reranking with Compressed Document Representation 21 May 2025 · 0 repositories · arXiv:2505.15394
-
RLBenchNet: The Right Network for the Right Reinforcement Learning Task 21 May 2025 · 1 repository · arXiv:2505.15040
-
Robo-DM: Data Management For Large Robot Datasets 21 May 2025 · 0 repositories · arXiv:2505.15558
-
SAMA-UNet: Enhancing Medical Image Segmentation with Self-Adaptive Mamba-Like Attention and Causal-Resonance Learning 21 May 2025 · 1 repository · arXiv:2505.15234
-
Scaling Diffusion Transformers Efficiently via μP 21 May 2025 · 1 repository · arXiv:2505.15270
-
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries 21 May 2025 · 1 repository · arXiv:2505.15420
-
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization 21 May 2025 · 0 repositories · arXiv:2505.15444
-
Small Language Models in the Real World: Insights from Industrial Text Classification 21 May 2025 · 0 repositories · arXiv:2505.16078
-
Sonnet: Spectral Operator Neural Network for Multivariable Time Series Forecasting 21 May 2025 · 1 repository · arXiv:2505.15312
-
AI-empowered Channel Estimation for Block-based Active IRS-enhanced Hybrid-field IoT Network 20 May 2025 · 0 repositories · arXiv:2505.14098
-
Articulatory Feature Prediction from Surface EMG during Speech Production 20 May 2025 · 1 repository · arXiv:2505.13814
-
Automatic Dataset Generation for Knowledge Intensive Question Answering Tasks 20 May 2025 · 0 repositories · arXiv:2505.14212
-
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation 20 May 2025 · 0 repositories · arXiv:2505.13957
-
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders 20 May 2025 · 0 repositories · arXiv:2505.14536Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code Generation 20 May 2025 · 1 repository · arXiv:2505.14646Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Choosing a Model, Shaping a Future: Comparing LLM Perspectives on Sustainability and its Relationship with AI 20 May 2025 · 0 repositories · arXiv:2505.14435
-
Cost-Augmented Monte Carlo Tree Search for LLM-Assisted Planning 20 May 2025 · 0 repositories · arXiv:2505.14656
-
Divide by Question, Conquer by Agent: SPLIT-RAG with Question-Driven Graph Partitioning 20 May 2025 · 0 repositories · arXiv:2505.13994
-
Do Language Models Use Their Depth Efficiently? 20 May 2025 · 1 repository · arXiv:2505.13898Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
DSMentor: Enhancing Data Science Agents with Curriculum Learning and Online Knowledge Accumulation 20 May 2025 · 0 repositories · arXiv:2505.14163
-
EEG-to-Text Translation: A Model for Deciphering Human Brain Activity 20 May 2025 · 1 repository · arXiv:2505.13936
-
Energy-Efficient Deep Reinforcement Learning with Spiking Transformers 20 May 2025 · 0 repositories · arXiv:2505.14533
-
Enhancing Abstractive Summarization of Scientific Papers Using Structure Information 20 May 2025 · 1 repository · arXiv:2505.14179
-
Every Pixel Tells a Story: End-to-End Urdu Newspaper OCR 20 May 2025 · 0 repositories · arXiv:2505.13943
-
FlashKAT: Understanding and Addressing Performance Bottlenecks in the Kolmogorov-Arnold Transformer 20 May 2025 · 1 repository · arXiv:2505.13813
-
Generative AI at the Crossroads: Light Bulb, Dynamo, or Microscope? 20 May 2025 · 0 repositories · arXiv:2505.14588
-
Informatics for Food Processing 20 May 2025 · 1 repository · arXiv:2505.17087
-
Latent Flow Transformer 20 May 2025 · 1 repository · arXiv:2505.14513
-
Learning Spatio-Temporal Dynamics for Trajectory Recovery via Time-Aware Transformer 20 May 2025 · 2 repositories · arXiv:2505.13857
-
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators 20 May 2025 · 0 repositories · arXiv:2505.14314
-
ModRWKV: Transformer Multimodality in Linear Time 20 May 2025 · 1 repository · arXiv:2505.14505
-
MSDformer: Multi-scale Discrete Transformer For Time Series Generation 20 May 2025 · 0 repositories · arXiv:2505.14202
-
Multi-Channel Swin Transformer Framework for Bearing Remaining Useful Life Prediction 20 May 2025 · 0 repositories · arXiv:2505.14897
-
Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models 20 May 2025 · 0 repositories · arXiv:2505.13828
-
OmniStyle: Filtering High Quality Style Transfer Data at Scale 20 May 2025 · 1 repository · arXiv:2505.14028
-
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective 20 May 2025 · 0 repositories · arXiv:2505.14808
-
Probing BERT for German Compound Semantics 20 May 2025 · 0 repositories · arXiv:2505.14130
-
Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning 20 May 2025 · 1 repository · arXiv:2505.14069Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ReactDiff: Latent Diffusion for Facial Reaction Generation 20 May 2025 · 1 repository · arXiv:2505.14151
-
s3: You Don't Need That Much Data to Train a Search Agent via RL 20 May 2025 · 1 repository · arXiv:2505.14146Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation 20 May 2025 · 1 repository · arXiv:2505.14821
-
Scaling Laws for State Dynamics in Large Language Models 20 May 2025 · 0 repositories · arXiv:2505.14892
-
SCAN: Semantic Document Layout Analysis for Textual and Visual Retrieval-Augmented Generation 20 May 2025 · 0 repositories · arXiv:2505.14381
-
Selective Structured State Space for Multispectral-fused Small Target Detection 20 May 2025 · 0 repositories · arXiv:2505.14043
-
Normalized Cut with Reinforcement Learning in Constrained Action Space 20 May 2025 · 0 repositories · arXiv:2505.13986
-
SSPS: Self-Supervised Positive Sampling for Robust Self-Supervised Speaker Verification 20 May 2025 · 1 repository · arXiv:2505.14561
-
STree: Speculative Tree Decoding for Hybrid State-Space Models 20 May 2025 · 0 repositories · arXiv:2505.14969
-
Subquadratic Algorithms and Hardness for Attention with Any Temperature 20 May 2025 · 0 repositories · arXiv:2505.14840
-
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis 20 May 2025 · 1 repository · arXiv:2505.14910
-
A3 : an Analytical Low-Rank Approximation Framework for Attention 19 May 2025 · 0 repositories · arXiv:2505.12942
-
Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps 19 May 2025 · 0 repositories · arXiv:2505.12731
-
Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities 19 May 2025 · 0 repositories · arXiv:2505.13195
-
AMAQA: A Metadata-based QA Dataset for RAG Systems 19 May 2025 · 0 repositories · arXiv:2505.13557
-
Are Large Language Models Good at Detecting Propaganda? 19 May 2025 · 0 repositories · arXiv:2505.13706
-
Climate Research Domain BERTs: Pretraining, Adaptation, and Evaluation 19 May 2025 · 0 repositories
-
CMLFormer: A Dual Decoder Transformer with Switching Point Learning for Code-Mixed Language Modeling 19 May 2025 · 0 repositories · arXiv:2505.12587
-
EAVIT: Efficient and Accurate Human Value Identification from Text data via LLMs 19 May 2025 · 0 repositories · arXiv:2505.12792
-
Effective and Transparent RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability 19 May 2025 · 1 repository · arXiv:2505.13258
-
Enhancing Latent Computation in Transformers with Latent Tokens 19 May 2025 · 0 repositories · arXiv:2505.12629
-
Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain 19 May 2025 · 0 repositories · arXiv:2505.13006
-
GuidedMorph: Two-Stage Deformable Registration for Breast MRI 19 May 2025 · 0 repositories · arXiv:2505.13414
-
Know Or Not: a library for evaluating out-of-knowledge base robustness 19 May 2025 · 1 repository · arXiv:2505.13545
-
Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering 19 May 2025 · 1 repository · arXiv:2505.12662
-
LiDAR MOT-DETR: A LiDAR-based Two-Stage Transformer for 3D Multiple Object Tracking 19 May 2025 · 0 repositories · arXiv:2505.12753
-
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion 19 May 2025 · 1 repository · arXiv:2505.14719
-
Multi-head Temporal Latent Attention 19 May 2025 · 1 repository · arXiv:2505.13544Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
OMGPT: A Sequence Modeling Framework for Data-driven Operational Decision Making 19 May 2025 · 0 repositories · arXiv:2505.13580
-
Optimizing Retrieval Augmented Generation for Object Constraint Language 19 May 2025 · 0 repositories · arXiv:2505.13129
-
PPTNet: A Hybrid Periodic Pattern-Transformer Architecture for Traffic Flow Prediction and Congestion Identification 19 May 2025 · 1 repository · arXiv:2505.13047
-
Pyramid Sparse Transformer: Enhancing Multi-Scale Feature Fusion with Dynamic Token Selection 19 May 2025 · 0 repositories · arXiv:2505.12772
-
RAR: Setting Knowledge Tripwires for Retrieval Augmented Rejection 19 May 2025 · 0 repositories · arXiv:2505.13581
-
Simplicity is Key: An Unsupervised Pretraining Approach for Sparse Radio Channels 19 May 2025 · 0 repositories · arXiv:2505.13055
-
SounDiT: Geo-Contextual Soundscape-to-Landscape Generation 19 May 2025 · 0 repositories · arXiv:2505.12734
-
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset 19 May 2025 · 0 repositories · arXiv:2505.13069
-
The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting 19 May 2025 · 0 repositories · arXiv:2505.12837