Methods › General › Attention Modules › Multi-Head Attention › Papers, page 23
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 23 of 249: papers 2,201 to 2,300 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Provence: efficient and robust context pruning for retrieval-augmented generation 27 Jan 2025 · 0 repositories · arXiv:2501.16214
-
RelCAT: Advancing Extraction of Clinical Inter-Entity Relationships from Unstructured Electronic Health Records 27 Jan 2025 · 1 repository · arXiv:2501.16077
-
UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-trained Speech Models for Robust Speaker Verification 27 Jan 2025 · 0 repositories · arXiv:2501.16542
-
URAG: Implementing a Unified Hybrid RAG for Precise Answers in University Admission Chatbots -- A Case Study at HCMUT 27 Jan 2025 · 0 repositories · arXiv:2501.16276
-
Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference 27 Jan 2025 · 0 repositories · arXiv:2501.15754
-
Adapting Biomedical Abstracts into Plain language using Large Language Models 26 Jan 2025 · 0 repositories · arXiv:2501.15700
-
AI-Driven Secure Data Sharing: A Trustworthy and Privacy-Preserving Approach 26 Jan 2025 · 0 repositories · arXiv:2501.15363
-
ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer 26 Jan 2025 · 1 repository · arXiv:2501.15570
-
Classifying Deepfakes Using Swin Transformers 26 Jan 2025 · 0 repositories · arXiv:2501.15656
-
Decentralized Low-Rank Fine-Tuning of Large Language Models 26 Jan 2025 · 0 repositories · arXiv:2501.15361
-
Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets 26 Jan 2025 · 0 repositories · arXiv:2501.15624
-
SedarEval: Automated Evaluation using Self-Adaptive Rubrics 26 Jan 2025 · 1 repository · arXiv:2501.15595Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets? 26 Jan 2025 · 0 repositories · arXiv:2501.15431
-
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel 26 Jan 2025 · 0 repositories · arXiv:2501.15665
-
TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs 26 Jan 2025 · 1 repository · arXiv:2501.15674
-
Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment 26 Jan 2025 · 0 repositories · arXiv:2501.16394
-
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics 26 Jan 2025 · 1 repository · arXiv:2501.17187
-
Advanced Real-Time Fraud Detection Using RAG-Based LLMs 25 Jan 2025 · 0 repositories · arXiv:2501.15290
-
An AI-Driven Live Systematic Reviews in the Brain-Heart Interconnectome: Minimizing Research Waste and Advancing Evidence Synthesis 25 Jan 2025 · 1 repository · arXiv:2501.17181
-
An Attempt to Unraveling Token Prediction Refinement and Identifying Essential Layers of Large Language Models 25 Jan 2025 · 0 repositories · arXiv:2501.15054
-
ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Retrieval 25 Jan 2025 · 0 repositories · arXiv:2501.15245
-
CG-RAG: Research Question Answering by Citation Graph Retrieval-Augmented LLMs 25 Jan 2025 · 0 repositories · arXiv:2501.15067
-
Crystal Oscillators in OSNMA-Enabled Receivers: An Implementation View for Automotive Applications 25 Jan 2025 · 0 repositories · arXiv:2501.15123
-
ILETIA: An AI-enhanced method for individualized trigger-oocyte pickup interval estimation of progestin-primed ovarian stimulation protocol 25 Jan 2025 · 0 repositories · arXiv:2501.16386
-
Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement Learning 25 Jan 2025 · 1 repository · arXiv:2501.15228Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Knowledge Hierarchy Guided Biological-Medical Dataset Distillation for Domain LLM Training 25 Jan 2025 · 0 repositories · arXiv:2501.15108
-
LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering 25 Jan 2025 · 0 repositories · arXiv:2501.17183
-
Speech Translation Refinement using Large Language Models 25 Jan 2025 · 1 repository · arXiv:2501.15090
-
TranStable: Towards Robust Pixel-level Online Video Stabilization by Jointing Transformer and CNN 25 Jan 2025 · 0 repositories · arXiv:2501.15138
-
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information 25 Jan 2025 · 0 repositories · arXiv:2502.15714
-
Using Large Language Models for education managements in Vietnamese with low resources 25 Jan 2025 · 0 repositories · arXiv:2501.15022
-
Advances in Set Function Learning: A Survey of Techniques and Applications 24 Jan 2025 · 0 repositories · arXiv:2501.14991
-
Automatic detection and prediction of nAMD activity change in retinal OCT using Siamese networks and Wasserstein Distance for ordinality 24 Jan 2025 · 1 repository · arXiv:2501.14323
-
Causal Graphs Meet Thoughts: Enhancing Complex Reasoning in Graph-Augmented LLMs 24 Jan 2025 · 1 repository · arXiv:2501.14892
-
Chain-of-Retrieval Augmented Generation 24 Jan 2025 · 0 repositories · arXiv:2501.14342
-
Characteristic-Specific Partial Fine-Tuning for Efficient Emotion and Speaker Adaptation in Codec Language Text-to-Speech Models 24 Jan 2025 · 0 repositories · arXiv:2501.14273
-
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs 24 Jan 2025 · 0 repositories · arXiv:2501.18617
-
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning 24 Jan 2025 · 0 repositories · arXiv:2501.14680
-
Fast Think-on-Graph: Wider, Deeper and Faster Reasoning of Large Language Model on Knowledge Graph 24 Jan 2025 · 1 repository · arXiv:2501.14300
-
GraPPI: A Retrieve-Divide-Solve GraphRAG Framework for Large-scale Protein-protein Interaction Exploration 24 Jan 2025 · 1 repository · arXiv:2501.16382
-
Idiom Detection in Sorani Kurdish Texts 24 Jan 2025 · 0 repositories · arXiv:2501.14528
-
Iterative Feature Space Optimization through Incremental Adaptive Evaluation 24 Jan 2025 · 0 repositories · arXiv:2501.14889
-
Low-rank Prompt Interaction for Continual Vision-Language Retrieval 24 Jan 2025 · 1 repository · arXiv:2501.14369
-
On the locality bias and results in the Long Range Arena 24 Jan 2025 · 0 repositories · arXiv:2501.14850
-
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant 24 Jan 2025 · 0 repositories · arXiv:2501.17176
-
Rethinking Table Instruction Tuning 24 Jan 2025 · 1 repository · arXiv:2501.14693
-
Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation 24 Jan 2025 · 0 repositories · arXiv:2501.14679
-
Test-Time Code-Switching for Cross-lingual Aspect Sentiment Triplet Extraction 24 Jan 2025 · 0 repositories · arXiv:2501.14144
-
UDiTQC: U-Net-Style Diffusion Transformer for Quantum Circuit Synthesis 24 Jan 2025 · 0 repositories · arXiv:2501.16380
-
ZETA: Leveraging Z-order Curves for Efficient Top-k Attention 24 Jan 2025 · 0 repositories · arXiv:2501.14577
-
5G LDPC Linear Transformer for Channel Decoding 23 Jan 2025 · 1 repository · arXiv:2501.14102
-
A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification 23 Jan 2025 · 1 repository · arXiv:2501.13598
-
An Efficient Diffusion-based Non-Autoregressive Solver for Traveling Salesman Problem 23 Jan 2025 · 1 repository · arXiv:2501.13767Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
CAPRAG: A Large Language Model Solution for Customer Service and Automatic Reporting using Vector and Graph Retrieval-Augmented Generation 23 Jan 2025 · 0 repositories · arXiv:2501.13993
-
Document-Level Sentiment Analysis of Urdu Text Using Deep Learning Techniques 23 Jan 2025 · 0 repositories · arXiv:2501.17175
-
EgoHand: Ego-centric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMUs 23 Jan 2025 · 1 repository · arXiv:2501.13805
-
Enhancing Biomedical Relation Extraction with Directionality 23 Jan 2025 · 1 repository · arXiv:2501.14079
-
Ensuring Medical AI Safety: Explainable AI-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data 23 Jan 2025 · 1 repository · arXiv:2501.13818
-
FreEformer: Frequency Enhanced Transformer for Multivariate Time Series Forecasting 23 Jan 2025 · 1 repository · arXiv:2501.13989Syntology official (archive's flag): 7 ran · 7 ran (of which 7 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; every one of the 7 samples that ran constructed an object rather than computing a result (of 10 harvested samples) · 10 pointer-only (licence)
-
GraphRAG under Fire 23 Jan 2025 · 0 repositories · arXiv:2501.14050
-
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language 23 Jan 2025 · 0 repositories · arXiv:2501.14073
-
LLMs Can Plan Only If We Tell Them 23 Jan 2025 · 0 repositories · arXiv:2501.13545
-
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods 23 Jan 2025 · 0 repositories · arXiv:2501.13484
-
ME-CPT: Multi-Task Enhanced Cross-Temporal Point Transformer for Urban 3D Change Detection 23 Jan 2025 · 1 repository · arXiv:2501.14004
-
Multi-Level Attention and Contrastive Learning for Enhanced Text Classification with an Optimized Transformer 23 Jan 2025 · 0 repositories · arXiv:2501.13467
-
Polyhedra Encoding Transformers: Enhancing Diffusion MRI Analysis Beyond Voxel and Volumetric Embedding 23 Jan 2025 · 0 repositories · arXiv:2501.13352
-
Quantized Spike-driven Transformer 23 Jan 2025 · 1 repository · arXiv:2501.13492Syntology official (archive's flag): 11 ran · 13 ran (of which 9 constructed an object rather than computing a result; 13 with no instrument failure: 2 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 12 unverified (of 25 harvested samples) · 25 pointer-only (licence)
-
Question Answering on Patient Medical Records with Private Fine-Tuned LLMs 23 Jan 2025 · 0 repositories · arXiv:2501.13687
-
Retrievals Can Be Detrimental: A Contrastive Backdoor Attack Paradigm on Retrieval-Augmented Diffusion Models 23 Jan 2025 · 0 repositories · arXiv:2501.13340
-
RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation 23 Jan 2025 · 0 repositories · arXiv:2501.13726
-
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models 23 Jan 2025 · 0 repositories · arXiv:2501.13629
-
StreamingRAG: Real-time Contextual Retrieval and Generation Framework 23 Jan 2025 · 0 repositories · arXiv:2501.14101
-
Text-driven Online Action Detection 23 Jan 2025 · 1 repository · arXiv:2501.13518
-
Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning 23 Jan 2025 · 1 repository · arXiv:2501.13883Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home 22 Jan 2025 · 0 repositories · arXiv:2501.12835Syntology 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 5 pointer-only (licence)
-
Ehrenfeucht-Haussler Rank and Chain of Thought 22 Jan 2025 · 0 repositories · arXiv:2501.12997
-
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model 22 Jan 2025 · 0 repositories · arXiv:2501.12682
-
EvidenceMap: Learning Evidence Analysis to Unleash the Power of Small Language Models for Biomedical Question Answering 22 Jan 2025 · 0 repositories · arXiv:2501.12746
-
Exploring GPT's Ability as a Judge in Music Understanding 22 Jan 2025 · 1 repository · arXiv:2501.13261
-
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana 22 Jan 2025 · 0 repositories · arXiv:2501.12789
-
LiT: Delving into a Simplified Linear Diffusion Transformer for Image Generation 22 Jan 2025 · 0 repositories · arXiv:2501.12976
-
Multimodal AI on Wound Images and Clinical Notes for Home Patient Referral 22 Jan 2025 · 0 repositories · arXiv:2501.13247
-
RAG-Reward: Optimizing RAG with Reward Modeling and RLHF 22 Jan 2025 · 0 repositories · arXiv:2501.13264
-
Separated Inter/Intra-Modal Fusion Prompts for Compositional Zero-Shot Learning 22 Jan 2025 · 0 repositories · arXiv:2501.17171
-
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding 22 Jan 2025 · 1 repository · arXiv:2501.13200
-
T-Graphormer: Using Transformers for Spatiotemporal Forecasting 22 Jan 2025 · 1 repository · arXiv:2501.13274
-
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi 22 Jan 2025 · 0 repositories · arXiv:2501.12900
-
A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models 21 Jan 2025 · 1 repository · arXiv:2501.13958
-
Academic Case Reports Lack Diversity: Assessing the Presence and Diversity of Sociodemographic and Behavioral Factors related to Post COVID-19 Condition 21 Jan 2025 · 0 repositories · arXiv:2501.12538
-
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble 21 Jan 2025 · 1 repository · arXiv:2501.13964
-
ALoFTRAG: Automatic Local Fine Tuning for Retrieval Augmented Generation 21 Jan 2025 · 1 repository · arXiv:2501.11929
-
Assisting Mathematical Formalization with A Learning-based Premise Retriever 21 Jan 2025 · 1 repository · arXiv:2501.13959
-
Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration 21 Jan 2025 · 0 repositories · arXiv:2501.12332
-
Comparative Approaches to Sentiment Analysis Using Datasets in Major European and Arabic Languages 21 Jan 2025 · 0 repositories · arXiv:2501.12540
-
Continuous 3D Perception Model with Persistent State 21 Jan 2025 · 0 repositories · arXiv:2501.12387
-
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation 21 Jan 2025 · 0 repositories · arXiv:2501.12432
-
DLEN: Dual Branch of Transformer for Low-Light Image Enhancement in Dual Domains 21 Jan 2025 · 0 repositories · arXiv:2501.12235
-
Efficient Lung Ultrasound Severity Scoring Using Dedicated Feature Extractor 21 Jan 2025 · 1 repository · arXiv:2501.12524
-
Episodic Memories Generation and Evaluation Benchmark for Large Language Models 21 Jan 2025 · 1 repository · arXiv:2501.13121Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
FOCUS: First Order Concentrated Updating Scheme 21 Jan 2025 · 0 repositories · arXiv:2501.12243