Methods › General › Attention Modules › Multi-Head Attention › Papers, page 22
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 22 of 249: papers 2,101 to 2,200 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models 3 Feb 2025 · 0 repositories · arXiv:2502.01386
-
Toward Neurosymbolic Program Comprehension 3 Feb 2025 · 0 repositories · arXiv:2502.01806
-
Transformers trained on proteins can learn to attend to Euclidean distance 3 Feb 2025 · 1 repository · arXiv:2502.01533
-
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos 3 Feb 2025 · 1 repository · arXiv:2502.01549
-
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale 2 Feb 2025 · 1 repository · arXiv:2502.01681Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Estimating forest carbon stocks from high-resolution remote sensing imagery by reducing domain shift with style transfer 2 Feb 2025 · 0 repositories · arXiv:2502.00784
-
Explainability in Practice: A Survey of Explainable NLP Across Various Domains 2 Feb 2025 · 0 repositories · arXiv:2502.00837
-
LIBRA: Measuring Bias of Large Language Model from a Local Context 2 Feb 2025 · 0 repositories · arXiv:2502.01679
-
A framework for river connectivity classification using temporal image processing and attention based neural networks 1 Feb 2025 · 0 repositories · arXiv:2502.00474
-
A Study on the Performance of U-Net Modifications in Retroperitoneal Tumor Segmentation 1 Feb 2025 · 1 repository · arXiv:2502.00314
-
Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset 1 Feb 2025 · 0 repositories · arXiv:2502.01676
-
CoddLLM: Empowering Large Language Models for Data Analytics 1 Feb 2025 · 0 repositories · arXiv:2502.00329
-
Contrastive Forward-Forward: A Training Algorithm of Vision Transformer 1 Feb 2025 · 0 repositories · arXiv:2502.00571
-
Converting Transformers into DGNNs Form 1 Feb 2025 · 1 repository · arXiv:2502.00585
-
Explainable AI for Sentiment Analysis of Human Metapneumovirus (HMPV) Using XLNet 1 Feb 2025 · 0 repositories · arXiv:2502.01663
-
Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms 1 Feb 2025 · 0 repositories · arXiv:2502.00234
-
MambaGlue: Fast and Robust Local Feature Matching With Mamba 1 Feb 2025 · 1 repository · arXiv:2502.00462
-
Milmer: a Framework for Multiple Instance Learning based Multimodal Emotion Recognition 1 Feb 2025 · 1 repository · arXiv:2502.00547
-
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation 1 Feb 2025 · 1 repository · arXiv:2502.00306
-
Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective 1 Feb 2025 · 0 repositories · arXiv:2502.00281
-
VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility 1 Feb 2025 · 1 repository · arXiv:2502.00543
-
Accelerating Diffusion Transformer via Error-Optimized Cache 31 Jan 2025 · 0 repositories · arXiv:2501.19243
-
Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers 31 Jan 2025 · 0 repositories · arXiv:2502.00070
-
CerraData-4MM: A multimodal benchmark dataset on Cerrado for land use and land cover classification 31 Jan 2025 · 1 repository · arXiv:2502.00083
-
ContextFormer: Redefining Efficiency in Semantic Segmentation 31 Jan 2025 · 0 repositories · arXiv:2501.19255
-
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game 31 Jan 2025 · 1 repository · arXiv:2501.19398
-
From Semantic Segmentation of Natural Images to Medical Image Segmentation Using ViT-Based Architectures 31 Jan 2025 · 0 repositories
-
Homogeneity Bias as Differential Sampling Uncertainty in Language Models 31 Jan 2025 · 0 repositories · arXiv:2501.19337
-
KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search 31 Jan 2025 · 1 repository · arXiv:2501.18922Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 8 harvested samples)
-
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media 31 Jan 2025 · 0 repositories · arXiv:2502.01658
-
PixelWorld: Towards Perceiving Everything as Pixels 31 Jan 2025 · 0 repositories · arXiv:2501.19339
-
Privacy Preserving Charge Location Prediction for Electric Vehicles 31 Jan 2025 · 0 repositories · arXiv:2502.00068
-
Strassen Attention: Unlocking Compositional Abilities in Transformers Based on a New Lower Bound Method 31 Jan 2025 · 0 repositories · arXiv:2501.19215
-
Through the Looking Glass: LLM-Based Analysis of AR/VR Android Applications Privacy Policies 31 Jan 2025 · 0 repositories · arXiv:2501.19223
-
A Learnable Multi-views Contrastive Framework with Reconstruction Discrepancy for Medical Time-Series 30 Jan 2025 · 0 repositories · arXiv:2501.18367
-
A Unified Perspective on the Dynamics of Deep Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18322
-
AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates 30 Jan 2025 · 0 repositories · arXiv:2501.18094
-
Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18237
-
Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method 30 Jan 2025 · 0 repositories · arXiv:2501.18539
-
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights 30 Jan 2025 · 0 repositories · arXiv:2501.18596
-
Economic Rationality under Specialization: Evidence of Decision Bias in AI Agents 30 Jan 2025 · 0 repositories · arXiv:2501.18190
-
Evaluating Large Language Models in Vulnerability Detection Under Variable Context Windows 30 Jan 2025 · 0 repositories · arXiv:2502.00064
-
GDformer: Going Beyond Subsequence Isolation for Multivariate Time Series Anomaly Detection 30 Jan 2025 · 1 repository · arXiv:2501.18196
-
General Embedding vs. Task-Specific Embedding: A Comparative Approach to Enhancing NLP Performance 30 Jan 2025 · 0 repositories
-
GENIE: Generative Note Information Extraction model for structuring EHR data 30 Jan 2025 · 0 repositories · arXiv:2501.18435
-
On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning 30 Jan 2025 · 0 repositories · arXiv:2501.18187
-
Israel-Hamas war through Telegram, Reddit and Twitter 30 Jan 2025 · 0 repositories · arXiv:2502.00060
-
Leveraging LLM Agents for Automated Optimization Modeling for SASP Problems: A Graph-RAG based Approach 30 Jan 2025 · 0 repositories · arXiv:2501.18320
-
MatIR: A Hybrid Mamba-Transformer Image Restoration Model 30 Jan 2025 · 1 repository · arXiv:2501.18401
-
RbFT: Robust Fine-tuning for Retrieval-Augmented Generation against Retrieval Defects 30 Jan 2025 · 1 repository · arXiv:2501.18365
-
Retrieval Augmented Generation Based LLM Evaluation For Protocol State Machine Inference With Chain-of-Thought Reasoning 30 Jan 2025 · 0 repositories · arXiv:2502.15727
-
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer 30 Jan 2025 · 1 repository · arXiv:2501.18427
-
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions 30 Jan 2025 · 0 repositories · arXiv:2502.12017
-
State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence 30 Jan 2025 · 0 repositories · arXiv:2501.18356
-
Structure Development in List-Sorting Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18666
-
Survey and Improvement Strategies for Gene Prioritization with Large Language Models 30 Jan 2025 · 0 repositories · arXiv:2501.18794
-
Transformer Semantic Genetic Programming for Symbolic Regression 30 Jan 2025 · 0 repositories · arXiv:2501.18479
-
Unraveling the Capabilities of Language Models in News Summarization 30 Jan 2025 · 1 repository · arXiv:2501.18128
-
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 30 Jan 2025 · 1 repository · arXiv:2501.18511Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
2SSP: A Two-Stage Framework for Structured Pruning of LLMs 29 Jan 2025 · 1 repository · arXiv:2501.17771
-
ContourFormer:Real-Time Contour-Based End-to-End Instance Segmentation Transformer 29 Jan 2025 · 1 repository · arXiv:2501.17688
-
DINT Transformer 29 Jan 2025 · 0 repositories · arXiv:2501.17486
-
Hybrid Graphs for Table-and-Text based Question Answering using LLMs 29 Jan 2025 · 0 repositories · arXiv:2501.17767
-
Leveraging In-Context Learning and Retrieval-Augmented Generation for Automatic Question Generation in Educational Domains 29 Jan 2025 · 0 repositories · arXiv:2501.17397
-
Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Model: An Initial Multi-layered Tabular Review 29 Jan 2025 · 0 repositories · arXiv:2501.18644
-
PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion 29 Jan 2025 · 1 repository · arXiv:2501.17699
-
Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling 29 Jan 2025 · 1 repository · arXiv:2501.17772
-
Shared DIFF Transformer 29 Jan 2025 · 0 repositories · arXiv:2501.17900
-
Transformer Based Time-Series Forecasting for Stock 29 Jan 2025 · 0 repositories · arXiv:2502.09625
-
TransRAD: Retentive Vision Transformer for Enhanced Radar Object Detection 29 Jan 2025 · 1 repository · arXiv:2501.17977
-
Watch Your STEPP: Semantic Traversability Estimation using Pose Projected Features 29 Jan 2025 · 0 repositories · arXiv:2501.17594
-
An Attention-Locating Algorithm for Eliminating Background Effects in Fine-grained Visual Classification 28 Jan 2025 · 1 repository
-
Attribution analysis of legal language as used by LLM 28 Jan 2025 · 0 repositories · arXiv:2501.17330
-
Balancing Content Size in RAG-Text2SQL System 28 Jan 2025 · 0 repositories · arXiv:2502.15723
-
Chinese Stock Prediction Based on a Multi-Modal Transformer Framework: Macro-Micro Information Fusion 28 Jan 2025 · 0 repositories · arXiv:2501.16621
-
Detecting harassment and defamation in cyberbullying with emotion-adaptive training 28 Jan 2025 · 1 repository · arXiv:2501.16925
-
FlexMotion: Lightweight, Physics-Aware, and Controllable Human Motion Generation 28 Jan 2025 · 0 repositories · arXiv:2501.16778
-
Generative quantum combinatorial optimization by means of a novel conditional generative quantum eigensolver 28 Jan 2025 · 0 repositories · arXiv:2501.16986
-
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation 28 Jan 2025 · 1 repository · arXiv:2501.18638
-
Graph Transformers for inverse physics: reconstructing flows around arbitrary 2D airfoils 28 Jan 2025 · 0 repositories · arXiv:2501.17081
-
JRE-L: Journalist, Reader, and Editor LLMs in the Loop for Science Journalism for the General Audience 28 Jan 2025 · 1 repository · arXiv:2501.16865
-
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models 28 Jan 2025 · 1 repository · arXiv:2501.17088Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition 28 Jan 2025 · 0 repositories · arXiv:2501.17011
-
Multiple Abstraction Level Retrieve Augment Generation 28 Jan 2025 · 0 repositories · arXiv:2501.16952
-
Open-Source Retrieval Augmented Generation Framework for Retrieving Accurate Medication Insights from Formularies for African Healthcare Workers 28 Jan 2025 · 0 repositories · arXiv:2502.15722
-
SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model 28 Jan 2025 · 1 repository · arXiv:2501.18636
-
Scenario Understanding of Traffic Scenes Through Large Visual Language Models 28 Jan 2025 · 0 repositories · arXiv:2501.17131
-
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification 28 Jan 2025 · 1 repository · arXiv:2501.17260
-
Multimodal Magic Elevating Depression Detection with a Fusion of Text and Audio Intelligence 28 Jan 2025 · 0 repositories · arXiv:2501.16813
-
A Comprehensive Study on Fine-Tuning Large Language Models for Medical Question Answering Using Classification Models and Comparative Analysis 27 Jan 2025 · 0 repositories · arXiv:2501.17190
-
Cross-Domain Semantic Segmentation with Large Language Model-Assisted Descriptor Generation 27 Jan 2025 · 0 repositories · arXiv:2501.16467
-
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0 27 Jan 2025 · 0 repositories · arXiv:2501.16201
-
Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice 27 Jan 2025 · 0 repositories · arXiv:2502.07088
-
LCTG Bench: LLM Controlled Text Generation Benchmark 27 Jan 2025 · 1 repository · arXiv:2501.15875
-
LemmaHead: RAG Assisted Proof Generation Using Large Language Models 27 Jan 2025 · 0 repositories · arXiv:2501.15797
-
Leveraging Video Vision Transformer for Alzheimer's Disease Diagnosis from 3D Brain MRI 27 Jan 2025 · 0 repositories · arXiv:2501.15733
-
Object Detection for Medical Image Analysis: Insights from the RT-DETR Model 27 Jan 2025 · 0 repositories · arXiv:2501.16469
-
Parametric Retrieval Augmented Generation 27 Jan 2025 · 1 repository · arXiv:2501.15915
-
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer 27 Jan 2025 · 0 repositories · arXiv:2501.16227
-
Phase Transitions in Large Language Models and the O(N) Model 27 Jan 2025 · 0 repositories · arXiv:2501.16241