Methods › General › Attention Modules › Multi-Head Attention › Papers, page 2
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 2 of 249: papers 101 to 200 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Machine vs Machine: Using AI to Tackle Generative AI Threats in Assessment 31 May 2025 · 0 repositories · arXiv:2506.02046
-
Power-of-Two (PoT) Weights in Large Language Models (LLMs) 31 May 2025 · 0 repositories · arXiv:2506.00315
-
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations 31 May 2025 · 1 repository · arXiv:2506.00748
-
Adversarial Threat Vectors and Risk Mitigation for Retrieval-Augmented Generation Systems 30 May 2025 · 0 repositories · arXiv:2506.00281
-
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks 30 May 2025 · 1 repository · arXiv:2505.24876
-
ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation 30 May 2025 · 1 repository · arXiv:2505.24388
-
Cross-Attention Speculative Decoding 30 May 2025 · 0 repositories · arXiv:2505.24544
-
D2AF: A Dual-Driven Annotation and Filtering Framework for Visual Grounding 30 May 2025 · 0 repositories · arXiv:2505.24372
-
Interpretable phenotyping of Heart Failure patients with Dutch discharge letters 30 May 2025 · 0 repositories · arXiv:2505.24619
-
Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning 30 May 2025 · 1 repository · arXiv:2505.24360
-
Leveraging Intermediate Features of Vision Transformer for Face Anti-Spoofing 30 May 2025 · 0 repositories · arXiv:2505.24402
-
LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs 30 May 2025 · 0 repositories · arXiv:2505.24451
-
Mamba Knockout for Unraveling Factual Information Flow 30 May 2025 · 1 repository · arXiv:2505.24244Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer 30 May 2025 · 1 repository · arXiv:2505.24378
-
Model-Guided Network with Cluster-Based Operators for Spatio-Spectral Super-Resolution 30 May 2025 · 1 repository · arXiv:2505.24605
-
MOFGPT: Generative Design of Metal-Organic Frameworks using Language Models 30 May 2025 · 1 repository · arXiv:2506.00198
-
PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge 30 May 2025 · 0 repositories · arXiv:2505.24411
-
PersianMedQA: Language-Centric Evaluation of LLMs in the Persian Medical Domain 30 May 2025 · 0 repositories · arXiv:2506.00250
-
RealDrive: Retrieval-Augmented Driving with Diffusion Models 30 May 2025 · 0 repositories · arXiv:2505.24808
-
SPPSFormer: High-quality Superpoint-based Transformer for Roof Plane Instance Segmentation from Point Clouds 30 May 2025 · 0 repositories · arXiv:2505.24475
-
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs 30 May 2025 · 0 repositories · arXiv:2506.00197
-
Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition 29 May 2025 · 2 repositories · arXiv:2505.23313
-
ATLAS: Learning to Optimally Memorize the Context at Test Time 29 May 2025 · 0 repositories · arXiv:2505.23735
-
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time 29 May 2025 · 0 repositories · arXiv:2505.23729
-
Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning 29 May 2025 · 0 repositories · arXiv:2505.23298
-
CF-DETR: Coarse-to-Fine Transformer for Real-Time Object Detection 29 May 2025 · 0 repositories · arXiv:2505.23317
-
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training 29 May 2025 · 0 repositories · arXiv:2505.23971
-
DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers 29 May 2025 · 1 repository · arXiv:2505.23694Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples)
-
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs 29 May 2025 · 0 repositories · arXiv:2505.23299
-
DATD3: Depthwise Attention Twin Delayed Deep Deterministic Policy Gradient For Model Free Reinforcement Learning Under Output Feedback Control 29 May 2025 · 0 repositories · arXiv:2505.23857
-
Daunce: Data Attribution through Uncertainty Estimation 29 May 2025 · 0 repositories · arXiv:2505.23223
-
Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking 29 May 2025 · 0 repositories · arXiv:2505.23117
-
Deep Modeling and Optimization of Medical Image Classification 29 May 2025 · 1 repository · arXiv:2505.23040
-
Differential Gated Self-Attention 29 May 2025 · 0 repositories · arXiv:2505.24054
-
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models 29 May 2025 · 0 repositories · arXiv:2505.24025
-
Enhancing LLM-Based Code Generation with Complexity Metrics: A Feedback-Driven Approach 29 May 2025 · 0 repositories · arXiv:2505.23953
-
Equivariant Spherical Transformer for Efficient Molecular Modeling 29 May 2025 · 0 repositories · arXiv:2505.23086
-
Evaluating AI capabilities in detecting conspiracy theories on YouTube 29 May 2025 · 1 repository · arXiv:2505.23570
-
From Images to Signals: Are Large Vision Models Useful for Time Series Analysis? 29 May 2025 · 0 repositories · arXiv:2505.24030
-
HyperPointFormer: Multimodal Fusion in 3D Space with Dual-Branch Cross-Attention Transformers 29 May 2025 · 1 repository · arXiv:2505.23206
-
Learning to Regulate: A New Event-Level Dataset of Capital Control Measures 29 May 2025 · 0 repositories · arXiv:2505.23025
-
Matryoshka Model Learning for Improved Elastic Student Models 29 May 2025 · 0 repositories · arXiv:2505.23337
-
MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment 29 May 2025 · 0 repositories · arXiv:2505.23634
-
Patient Domain Supervised Contrastive Learning for Lung Sound Classification Using Mobile Phone 29 May 2025 · 0 repositories · arXiv:2505.23132
-
Probing Association Biases in LLM Moderation Over-Sensitivity 29 May 2025 · 0 repositories · arXiv:2505.23914
-
Query Routing for Retrieval-Augmented Language Models 29 May 2025 · 0 repositories · arXiv:2505.23052
-
Reducing Latency in LLM-Based Natural Language Commands Processing for Robot Navigation 29 May 2025 · 0 repositories · arXiv:2506.00075
-
Table-R1: Inference-Time Scaling for Table Reasoning 29 May 2025 · 1 repository · arXiv:2505.23621
-
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence 29 May 2025 · 1 repository · arXiv:2505.23420
-
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos 29 May 2025 · 1 repository · arXiv:2505.23693
-
Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems 28 May 2025 · 0 repositories · arXiv:2505.22571
-
Attention-Enhanced Prompt Decision Transformers for UAV-Assisted Communications with AoI 28 May 2025 · 0 repositories · arXiv:2505.22170
-
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon 28 May 2025 · 0 repositories · arXiv:2505.22184
-
Climate Finance Bench 28 May 2025 · 1 repository · arXiv:2505.22752
-
Contextual Memory Intelligence -- A Foundational Paradigm for Human-AI Collaboration and Reflective Generative AI Systems 28 May 2025 · 0 repositories · arXiv:2506.05370
-
Cross-modal RAG: Sub-dimensional Retrieval-Augmented Text-to-Image Generation 28 May 2025 · 1 repository · arXiv:2505.21956
-
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer 28 May 2025 · 2 repositories · arXiv:2505.22705Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs 28 May 2025 · 0 repositories · arXiv:2505.22937
-
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection 28 May 2025 · 0 repositories · arXiv:2505.22517
-
MultiFormer: A Multi-Person Pose Estimation System Based on CSI and Attention Mechanism 28 May 2025 · 0 repositories · arXiv:2505.22555
-
RAGPPI: RAG Benchmark for Protein-Protein Interactions in Drug Discovery 28 May 2025 · 1 repository · arXiv:2505.23823
-
Say What You Mean: Natural Language Access Control with Large Language Models for Internet of Things 28 May 2025 · 0 repositories · arXiv:2505.23835
-
SkewRoute: Training-Free LLM Routing for Knowledge Graph Retrieval-Augmented Generation via Score Skewness of Retrieved Context 28 May 2025 · 0 repositories · arXiv:2505.23841
-
UP-SLAM: Adaptively Structured Gaussian SLAM with Uncertainty Prediction in Dynamic Environments 28 May 2025 · 0 repositories · arXiv:2505.22335
-
Update Your Transformer to the Latest Release: Re-Basin of Task Vectors 28 May 2025 · 1 repository · arXiv:2505.22697Syntology official (archive's flag): 1 ran · 5 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 3 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning 28 May 2025 · 1 repository · arXiv:2505.22019Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A domain adaptation neural network for digital twin-supported fault diagnosis 27 May 2025 · 1 repository · arXiv:2505.21046
-
AgriFM: A Multi-source Temporal Remote Sensing Foundation Model for Crop Mapping 27 May 2025 · 1 repository · arXiv:2505.21357
-
Beyond 1D: Vision Transformers and Multichannel Signal Images for PPG-to-ECG Reconstruction 27 May 2025 · 0 repositories · arXiv:2505.21767
-
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers 27 May 2025 · 0 repositories · arXiv:2505.20666
-
Diagnosing and Resolving Cloud Platform Instability with Multi-modal RAG LLMs 27 May 2025 · 0 repositories · arXiv:2505.21419
-
Explainability of Large Language Models using SMILE: Statistical Model-agnostic Interpretability with Local Explanations 27 May 2025 · 1 repository · arXiv:2505.21657
-
From prosthetic memory to prosthetic denial: Auditing whether large language models are prone to mass atrocity denialism 27 May 2025 · 0 repositories · arXiv:2505.21753
-
HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling 27 May 2025 · 0 repositories · arXiv:2505.20836
-
HTMNet: A Hybrid Network with Transformer-Mamba Bottleneck Multimodal Fusion for Transparent and Reflective Objects Depth Completion 27 May 2025 · 0 repositories · arXiv:2505.20904
-
Long Context Scaling: Divide and Conquer via Multi-Agent Question-driven Collaboration 27 May 2025 · 0 repositories · arXiv:2505.20625
-
Minute-Long Videos with Dual Parallelisms 27 May 2025 · 1 repository · arXiv:2505.21070
-
MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition 27 May 2025 · 0 repositories · arXiv:2505.20744
-
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers 27 May 2025 · 0 repositories · arXiv:2505.21024
-
Privacy-Preserving Chest X-ray Report Generation via Multimodal Federated Learning with ViT and GPT-2 27 May 2025 · 0 repositories · arXiv:2505.21715
-
SOSBENCH: Benchmarking Safety Alignment on Scientific Knowledge 27 May 2025 · 0 repositories · arXiv:2505.21605
-
Absolute Coordinates Make Motion Generation Easy 26 May 2025 · 0 repositories · arXiv:2505.19377
-
Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation 26 May 2025 · 0 repositories · arXiv:2505.19554
-
AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and Healthcare 26 May 2025 · 1 repository · arXiv:2505.19562
-
Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents 26 May 2025 · 0 repositories · arXiv:2505.19494
-
Automated evaluation of children's speech fluency for low-resource languages 26 May 2025 · 0 repositories · arXiv:2505.19671
-
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models 26 May 2025 · 1 repository · arXiv:2505.19509
-
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages 26 May 2025 · 0 repositories · arXiv:2505.19851
-
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement 26 May 2025 · 1 repository · arXiv:2505.19675
-
CardioPatternFormer: Pattern-Guided Attention for Interpretable ECG Classification with Transformer Architecture 26 May 2025 · 0 repositories · arXiv:2505.20481
-
Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language 26 May 2025 · 0 repositories · arXiv:2505.19971
-
Dependency Parsing is More Parameter-Efficient with Normalization 26 May 2025 · 0 repositories · arXiv:2505.20215
-
Detection of Suicidal Risk on Social Media: A Hybrid Model 26 May 2025 · 0 repositories · arXiv:2505.23797
-
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems 26 May 2025 · 0 repositories · arXiv:2505.19847
-
DoctorRAG: Medical RAG Fusing Knowledge with Patient Analogy through Textual Gradients 26 May 2025 · 0 repositories · arXiv:2505.19538
-
Electrolyzers-HSI: Close-Range Multi-Scene Hyperspectral Imaging Benchmark Dataset 26 May 2025 · 0 repositories · arXiv:2505.20507
-
Emotion Classification In-Context in Spanish 26 May 2025 · 0 repositories · arXiv:2505.20571
-
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining 26 May 2025 · 0 repositories · arXiv:2505.19893
-
GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis 26 May 2025 · 1 repository · arXiv:2505.19813Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 6 where Syntology's instrument failed) · 9 unverified (of 24 harvested samples) · 24 pointer-only (licence)
-
Grokking ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior 26 May 2025 · 1 repository · arXiv:2505.20076Syntology official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 5 harvested samples)