Methods › General › Attention Modules › Multi-Head Attention › Papers, page 16
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 16 of 249: papers 1,501 to 1,600 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Evaluating Large Language Models in Code Generation: INFINITE Methodology for Defining the Inference Index 7 Mar 2025 · 0 repositories · arXiv:2503.05852
-
Simulating and Analysing Human Survey Responses with Large Language Models: A Case Study in Energy Stated Preference 7 Mar 2025 · 0 repositories · arXiv:2503.10652
-
Evaluating open-source Large Language Models for automated fact-checking 7 Mar 2025 · 0 repositories · arXiv:2503.05565
-
FastMap: Fast Queries Initialization Based Vectorized HD Map Reconstruction Framework 7 Mar 2025 · 1 repository · arXiv:2503.05492
-
FMCHS: Advancing Traditional Chinese Medicine Herb Recommendation with Fusion of Multiscale Correlations of Herbs and Symptoms 7 Mar 2025 · 0 repositories · arXiv:2503.05167
-
FMT:A Multimodal Pneumonia Detection Model Based on Stacking MOE Framework 7 Mar 2025 · 0 repositories · arXiv:2503.05626
-
Language modelling techniques for analysing the impact of human genetic variation 7 Mar 2025 · 0 repositories · arXiv:2503.10655
-
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation 7 Mar 2025 · 0 repositories · arXiv:2503.05530
-
Leveraging Semantic Type Dependencies for Clinical Named Entity Recognition 7 Mar 2025 · 0 repositories · arXiv:2503.05373
-
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice 7 Mar 2025 · 0 repositories · arXiv:2503.05978
-
Quantifying the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data 7 Mar 2025 · 0 repositories · arXiv:2503.05587
-
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning 7 Mar 2025 · 5 repositories · arXiv:2503.05592Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Tractable Representations for Convergent Approximation of Distributional HJB Equations 7 Mar 2025 · 0 repositories · arXiv:2503.05563
-
Zero-shot Medical Event Prediction Using a Generative Pre-trained Transformer on Electronic Health Records 7 Mar 2025 · 0 repositories · arXiv:2503.05893
-
A Generalist Cross-Domain Molecular Learning Framework for Structure-Based Drug Discovery 6 Mar 2025 · 0 repositories · arXiv:2503.04362
-
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning 6 Mar 2025 · 0 repositories · arXiv:2503.04973
-
BicliqueEncoder: An Efficient Method for Link Prediction in Bipartite Networks using Formal Concept Analysis and Transformer Encoder 6 Mar 2025 · 0 repositories · arXiv:2503.07645
-
Can We Optimize Deep RL Policy Weights as Trajectory Modeling? 6 Mar 2025 · 0 repositories · arXiv:2503.04074
-
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence 6 Mar 2025 · 0 repositories · arXiv:2503.05037
-
Compositional Causal Reasoning Evaluation in Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04556
-
DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQL 6 Mar 2025 · 0 repositories · arXiv:2503.04959
-
GBT-SAM: Adapting a Foundational Deep Learning Model for Generalizable Brain Tumor Segmentation via Efficient Integration of Multi-Parametric MRI Data 6 Mar 2025 · 1 repository · arXiv:2503.04325
-
Hedging with Sparse Reward Reinforcement Learning 6 Mar 2025 · 0 repositories · arXiv:2503.04218
-
High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects 6 Mar 2025 · 0 repositories · arXiv:2503.04862
-
HILGEN: Hierarchically-Informed Data Generation for Biomedical NER Using Knowledgebases and Large Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04930
-
In-depth Analysis of Graph-based RAG in a Unified Framework 6 Mar 2025 · 0 repositories · arXiv:2503.04338
-
Incentivizing Multi-Tenant Split Federated Learning for Foundation Models at the Network Edge 6 Mar 2025 · 0 repositories · arXiv:2503.04971
-
Learning Transformer-based World Models with Contrastive Predictive Coding 6 Mar 2025 · 0 repositories · arXiv:2503.04416
-
Leveraging Large Language Models to Address Data Scarcity in Machine Learning: Applications in Graphene Synthesis 6 Mar 2025 · 1 repository · arXiv:2503.04870
-
Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation 6 Mar 2025 · 1 repository · arXiv:2503.04871
-
Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04280
-
A Multimodal Framework for Topic Propagation Classification in Social Networks 5 Mar 2025 · 0 repositories · arXiv:2503.03112
-
AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything Model 5 Mar 2025 · 0 repositories · arXiv:2503.03088
-
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials 5 Mar 2025 · 1 repository · arXiv:2503.03965Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Can Frontier LLMs Replace Annotators in Biomedical Text Mining? Analyzing Challenges and Exploring Solutions 5 Mar 2025 · 1 repository · arXiv:2503.03261
-
DTU-Net: A Multi-Scale Dilated Transformer Network for Nonlinear Hyperspectral Unmixing 5 Mar 2025 · 0 repositories · arXiv:2503.03465
-
Intermediate-Task Transfer Learning: Leveraging Sarcasm Detection for Stance Detection 5 Mar 2025 · 0 repositories · arXiv:2503.03172
-
Large language models in finance : what is financial sentiment? 5 Mar 2025 · 0 repositories · arXiv:2503.03612
-
MA-LoT: Multi-Agent Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving 5 Mar 2025 · 1 repository · arXiv:2503.03205
-
PathRWKV: Enabling Whole Slide Prediction with Recurrent-Transformer 5 Mar 2025 · 0 repositories · arXiv:2503.03199
-
Personalized Federated Fine-tuning for Heterogeneous Data: An Automatic Rank Learning Approach via Two-Level LoRA 5 Mar 2025 · 0 repositories · arXiv:2503.03920
-
Pretrained LLMs as Real-Time Controllers for Robot Operated Serial Production Line 5 Mar 2025 · 0 repositories · arXiv:2503.03889
-
RiskAgent: Autonomous Medical AI Copilot for Generalist Risk Prediction 5 Mar 2025 · 0 repositories · arXiv:2503.03802
-
Sarcasm Detection as a Catalyst: Improving Stance Detection with Cross-Target Capabilities 5 Mar 2025 · 0 repositories · arXiv:2503.03787
-
ScaleFusionNet: Transformer-Guided Multi-Scale Feature Fusion for Skin Lesion Segmentation 5 Mar 2025 · 1 repository · arXiv:2503.03327
-
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation 5 Mar 2025 · 1 repository · arXiv:2503.03308
-
A Transformer Model for Predicting Chemical Reaction Products from Generic Templates 4 Mar 2025 · 0 repositories · arXiv:2503.05810
-
BHViT: Binarized Hybrid Vision Transformer 4 Mar 2025 · 1 repository · arXiv:2503.02394Syntology official (archive's flag): 16 ran · 16 ran (of which 13 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 1 where Syntology's instrument failed) · 13 unverified (of 29 harvested samples)
-
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory 4 Mar 2025 · 0 repositories · arXiv:2503.02354
-
Developing a PET/CT Foundation Model for Cross-Modal Anatomical and Functional Imaging 4 Mar 2025 · 0 repositories · arXiv:2503.02824
-
Effectively Steer LLM To Follow Preference via Building Confident Directions 4 Mar 2025 · 0 repositories · arXiv:2503.02989
-
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation 4 Mar 2025 · 0 repositories · arXiv:2503.07630
-
Graph Transformer with Disease Subgraph Positional Encoding for Improved Comorbidity Prediction 4 Mar 2025 · 1 repository · arXiv:2503.03046
-
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models 4 Mar 2025 · 0 repositories · arXiv:2503.02917
-
Learning Precoding in Multi-user Multi-antenna Systems: Transformer or Graph Transformer? 4 Mar 2025 · 0 repositories · arXiv:2503.02998
-
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning 4 Mar 2025 · 0 repositories · arXiv:2503.04812
-
Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models 4 Mar 2025 · 0 repositories · arXiv:2503.02141
-
Optimizing open-domain question answering with graph-based retrieval augmented generation 4 Mar 2025 · 0 repositories · arXiv:2503.02922
-
PennyLang: Pioneering LLM-Based Quantum Code Generation with a Novel PennyLane-Centric Dataset 4 Mar 2025 · 0 repositories · arXiv:2503.02497
-
Tabby: Tabular Data Synthesis with Language Models 4 Mar 2025 · 0 repositories · arXiv:2503.02152
-
Target Return Optimizer for Multi-Game Decision Transformer 4 Mar 2025 · 0 repositories · arXiv:2503.02311
-
TeTRA-VPR: A Ternary Transformer Approach for Compact Visual Place Recognition 4 Mar 2025 · 0 repositories · arXiv:2503.02511
-
Union of Experts: Adapting Hierarchical Routing to Equivalently Decomposed Transformer 4 Mar 2025 · 1 repository · arXiv:2503.02495
-
Use Me Wisely: AI-Driven Assessment for LLM Prompting Skills Development 4 Mar 2025 · 0 repositories · arXiv:2503.02532
-
Weak-to-Strong Generalization Even in Random Feature Networks, Provably 4 Mar 2025 · 0 repositories · arXiv:2503.02877
-
Wikipedia in the Era of LLMs: Evolution and Risks 4 Mar 2025 · 1 repository · arXiv:2503.02879
-
Wyckoff Transformer: Generation of Symmetric Crystals 4 Mar 2025 · 1 repository · arXiv:2503.02407
-
Zero-Shot Multi-Label Classification of Bangla Documents: Large Decoders Vs. Classic Encoders 4 Mar 2025 · 0 repositories · arXiv:2503.02993
-
From Claims to Evidence: A Unified Framework and Critical Analysis of CNN vs. Transformer vs. Mamba in Medical Image Segmentation 3 Mar 2025 · 1 repository · arXiv:2503.01306
-
Enhancing Social Media Rumor Detection: A Semantic and Graph Neural Network Approach for the 2024 Global Election 3 Mar 2025 · 0 repositories · arXiv:2503.01394
-
SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement Learning 3 Mar 2025 · 0 repositories · arXiv:2503.01458
-
An Efficient Approach to Detecting Lung Nodules Using Swin Transformer 3 Mar 2025 · 0 repositories · arXiv:2503.01592
-
Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data 3 Mar 2025 · 0 repositories · arXiv:2503.01630
-
SAGE: A Framework of Precise Retrieval for RAG 3 Mar 2025 · 0 repositories · arXiv:2503.01713
-
LLMInit: A Free Lunch from Large Language Models for Selective Initialization of Recommendation 3 Mar 2025 · 0 repositories · arXiv:2503.01814
-
A Hybrid CNN-Transformer Model for Heart Disease Prediction Using Life History Data 3 Mar 2025 · 0 repositories · arXiv:2503.02124
-
Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling 3 Mar 2025 · 1 repository · arXiv:2503.01215
-
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification 3 Mar 2025 · 0 repositories · arXiv:2503.01940
-
Attention Condensation via Sparsity Induced Regularized Training 3 Mar 2025 · 0 repositories · arXiv:2503.01564
-
Boolean-aware Attention for Dense Retrieval 3 Mar 2025 · 0 repositories · arXiv:2503.01753
-
Cancer Type, Stage and Prognosis Assessment from Pathology Reports using LLMs 3 Mar 2025 · 1 repository · arXiv:2503.01194
-
Dementia Insights: A Context-Based MultiModal Approach 3 Mar 2025 · 0 repositories · arXiv:2503.01226
-
Efficient or Powerful? Trade-offs Between Machine Learning and Deep Learning for Mental Illness Detection on Social Media 3 Mar 2025 · 0 repositories · arXiv:2503.01082
-
Forgetting Transformer: Softmax Attention with a Forget Gate 3 Mar 2025 · 1 repository · arXiv:2503.02130Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
HeterRec: Heterogeneous Information Transformer for Scalable Sequential Recommendation 3 Mar 2025 · 0 repositories · arXiv:2503.01469
-
HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation 3 Mar 2025 · 0 repositories · arXiv:2503.04800Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
How simple can you go? An off-the-shelf transformer approach to molecular dynamics 3 Mar 2025 · 1 repository · arXiv:2503.01431
-
Interactive Gadolinium-Free MRI Synthesis: A Transformer with Localization Prompt Learning 3 Mar 2025 · 1 repository · arXiv:2503.01265
-
Label Ranker: Self-Aware Preference for Classification Label Position in Visual Masked Self-Supervised Pre-Trained Model 3 Mar 2025 · 1 repository
-
MeshPad: Interactive Sketch-Conditioned Artist-Designed Mesh Generation and Editing 3 Mar 2025 · 0 repositories · arXiv:2503.01425
-
MI-DETR: An Object Detection Model with Multi-time Inquiries Mechanism 3 Mar 2025 · 1 repository · arXiv:2503.01463
-
MRI super-resolution reconstruction using efficient diffusion probabilistic model with residual shifting 3 Mar 2025 · 1 repository · arXiv:2503.01576
-
Primus: Enforcing Attention Usage for 3D Medical Image Segmentation 3 Mar 2025 · 0 repositories · arXiv:2503.01835
-
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG 3 Mar 2025 · 1 repository · arXiv:2503.01222
-
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction 3 Mar 2025 · 1 repository · arXiv:2503.01478Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
SRAG: Structured Retrieval-Augmented Generation for Multi-Entity Question Answering over Wikipedia Graph 3 Mar 2025 · 0 repositories · arXiv:2503.01346
-
Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection 3 Mar 2025 · 0 repositories · arXiv:2503.01362
-
Syntactic Learnability of Echo State Neural Language Models at Scale 3 Mar 2025 · 0 repositories · arXiv:2503.01724
-
Unify and Anchor: A Context-Aware Transformer for Cross-Domain Time Series Forecasting 3 Mar 2025 · 0 repositories · arXiv:2503.01157
-
Using (Not so) Large Language Models for Generating Simulation Models in a Formal DSL -- A Study on Reaction Networks 3 Mar 2025 · 0 repositories · arXiv:2503.01675