Methods › General › Attention Mechanisms › Attention › Papers, page 8
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 8 of 316: papers 701 to 800 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training 22 May 2025 · 1 repository · arXiv:2505.16363Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Align-GRAG: Reasoning-Guided Dual Alignment for Graph Retrieval-Augmented Generation 22 May 2025 · 0 repositories · arXiv:2505.16237
-
ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation 22 May 2025 · 1 repository · arXiv:2505.16495
-
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer 22 May 2025 · 0 repositories · arXiv:2505.16463
-
Assessing the generalization performance of SAM for ureteroscopy scene understanding 22 May 2025 · 0 repositories · arXiv:2505.17210
-
Attention with Trained Embeddings Provably Selects Important Tokens 22 May 2025 · 0 repositories · arXiv:2505.17282
-
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation 22 May 2025 · 0 repositories · arXiv:2505.16415Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA 22 May 2025 · 0 repositories · arXiv:2505.16293
-
Backdoor Cleaning without External Guidance in MLLM Fine-tuning 22 May 2025 · 1 repository · arXiv:2505.16916Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Beamforming-Codebook-Aware Channel Knowledge Map Construction for Multi-Antenna Systems 22 May 2025 · 1 repository · arXiv:2505.16132
-
Better Rates for Private Linear Regression in the Proportional Regime via Aggressive Clipping 22 May 2025 · 0 repositories · arXiv:2505.16329
-
Bottlenecked Transformers: Periodic KV Cache Abstraction for Generalised Reasoning 22 May 2025 · 0 repositories · arXiv:2505.16950
-
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention 22 May 2025 · 0 repositories · arXiv:2505.16157
-
CAIFormer: A Causal Informed Transformer for Multivariate Time Series Forecasting 22 May 2025 · 0 repositories · arXiv:2505.16308
-
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems 22 May 2025 · 0 repositories · arXiv:2505.16367
-
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks 22 May 2025 · 0 repositories · arXiv:2505.16901
-
Creatively Upscaling Images with Global-Regional Priors 22 May 2025 · 0 repositories · arXiv:2505.16976
-
DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes 22 May 2025 · 0 repositories · arXiv:2505.17162
-
Data-Driven Breakthroughs and Future Directions in AI Infrastructure: A Comprehensive Review 22 May 2025 · 0 repositories · arXiv:2505.16771
-
Emotion-based Recommender System 22 May 2025 · 0 repositories · arXiv:2505.16121
-
Erased or Dormant? Rethinking Concept Erasure Through Reversibility 22 May 2025 · 0 repositories · arXiv:2505.16174
-
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning 22 May 2025 · 0 repositories · arXiv:2505.16227
-
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles 22 May 2025 · 0 repositories · arXiv:2505.16784
-
Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification 22 May 2025 · 0 repositories · arXiv:2505.16338
-
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis 22 May 2025 · 1 repository · arXiv:2505.17241
-
Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models 22 May 2025 · 1 repository · arXiv:2505.16104
-
Internal Bias in Reasoning Models leads to Overthinking 22 May 2025 · 0 repositories · arXiv:2505.16448
-
JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation Model 22 May 2025 · 1 repository · arXiv:2505.17257Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Joint Flow And Feature Refinement Using Attention For Video Restoration 22 May 2025 · 0 repositories · arXiv:2505.16434
-
Learning Normal Patterns in Musical Loops 22 May 2025 · 0 repositories · arXiv:2505.23784
-
LINEA: Fast and Accurate Line Detection Using Scalable Transformers 22 May 2025 · 1 repository · arXiv:2505.16264
-
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion 22 May 2025 · 0 repositories · arXiv:2505.16565
-
MAGE: A Multi-task Architecture for Gaze Estimation with an Efficient Calibration Module 22 May 2025 · 0 repositories · arXiv:2505.16384
-
Meta-reinforcement learning with minimum attention 22 May 2025 · 0 repositories · arXiv:2505.16741
-
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression 22 May 2025 · 1 repository · arXiv:2505.16411
-
Multimodal Generative AI for Story Point Estimation in Software Development 22 May 2025 · 0 repositories · arXiv:2505.16290
-
Native Segmentation Vision Transformers 22 May 2025 · 0 repositories · arXiv:2505.16993
-
Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse 22 May 2025 · 0 repositories · arXiv:2505.16284
-
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning 22 May 2025 · 1 repository · arXiv:2505.16974
-
Partitioning and Observability in Linear Systems via Submodular Optimization 22 May 2025 · 0 repositories · arXiv:2505.16169
-
PaTH Attention: Position Encoding via Accumulating Householder Transformations 22 May 2025 · 2 repositories · arXiv:2505.16381Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Personalizing Student-Agent Interactions Using Log-Contextualized Retrieval Augmented Generation (RAG) 22 May 2025 · 0 repositories · arXiv:2505.17238
-
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction 22 May 2025 · 0 repositories · arXiv:2505.16980
-
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 22 May 2025 · 3 repositories · arXiv:2505.17005Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
RAP: Runtime-Adaptive Pruning for LLM Inference 22 May 2025 · 0 repositories · arXiv:2505.17138
-
Realistic Evaluation of TabPFN v2 in Open Environments 22 May 2025 · 0 repositories · arXiv:2505.16226
-
Reasoning in Neurosymbolic AI 22 May 2025 · 0 repositories · arXiv:2505.20313
-
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training 22 May 2025 · 1 repository · arXiv:2505.16792Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval 22 May 2025 · 0 repositories · arXiv:2505.16756
-
Reward-Aware Proto-Representations in Reinforcement Learning 22 May 2025 · 0 repositories · arXiv:2505.16217
-
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning 22 May 2025 · 0 repositories · arXiv:2505.16186
-
SAMba-UNet: Synergizing SAM2 and Mamba in UNet with Heterogeneous Aggregation for Cardiac MRI Segmentation 22 May 2025 · 0 repositories · arXiv:2505.16304
-
Scalable Graph Generative Modeling via Substructure Sequences 22 May 2025 · 1 repository · arXiv:2505.16130Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty 22 May 2025 · 0 repositories · arXiv:2505.17281
-
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding 22 May 2025 · 0 repositories · arXiv:2505.16652
-
Self-Classification Enhancement and Correction for Weakly Supervised Object Detection 22 May 2025 · 0 repositories · arXiv:2505.16294
-
SELF: Self-Extend the Context Length With Logistic Growth Function 22 May 2025 · 1 repository · arXiv:2505.17296
-
SHaDe: Compact and Consistent Dynamic 3D Reconstruction via Tri-Plane Deformation and Latent Diffusion 22 May 2025 · 0 repositories · arXiv:2505.16535
-
Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs 22 May 2025 · 0 repositories · arXiv:2505.16894
-
Sketchy Bounding-box Supervision for 3D Instance Segmentation 22 May 2025 · 1 repository · arXiv:2505.16399
-
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet 22 May 2025 · 0 repositories · arXiv:2505.16195
-
Style Transfer with Diffusion Models for Synthetic-to-Real Domain Adaptation 22 May 2025 · 1 repository · arXiv:2505.16360
-
Swin Transformer for Robust CGI Images Detection: Intra- and Inter-Dataset Analysis across Multiple Color Spaces 22 May 2025 · 0 repositories · arXiv:2505.16253
-
Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition 22 May 2025 · 0 repositories · arXiv:2505.16372
-
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks 22 May 2025 · 0 repositories · arXiv:2505.16594
-
The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm 22 May 2025 · 0 repositories · arXiv:2505.16932
-
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers 22 May 2025 · 0 repositories · arXiv:2505.16241
-
Training-Free Efficient Video Generation via Dynamic Token Carving 22 May 2025 · 1 repository · arXiv:2505.16864
-
Training-Free Reasoning and Reflection in MLLMs 22 May 2025 · 0 repositories · arXiv:2505.16151
-
Transformer brain encoders explain human high-level visual responses 22 May 2025 · 1 repository · arXiv:2505.17329
-
Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning 22 May 2025 · 1 repository · arXiv:2505.16270Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms 22 May 2025 · 0 repositories · arXiv:2505.17190Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Understanding Differential Transformer Unchains Pretrained Self-Attentions 22 May 2025 · 0 repositories · arXiv:2505.16333
-
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering 22 May 2025 · 0 repositories · arXiv:2505.17326
-
Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks 22 May 2025 · 1 repository · arXiv:2505.16849
-
When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction 22 May 2025 · 1 repository · arXiv:2505.16170
-
Why Can Accurate Models Be Learned from Inaccurate Annotations? 22 May 2025 · 0 repositories · arXiv:2505.16159
-
Zebra-Llama: Towards Extremely Efficient Hybrid Models 22 May 2025 · 0 repositories · arXiv:2505.17272
-
Zero-Shot Hyperspectral Pansharpening Using Hysteresis-Based Tuning for Spectral Quality Control 22 May 2025 · 1 repository · arXiv:2505.16658
-
A Taxonomy of Structure from Motion Methods 21 May 2025 · 0 repositories · arXiv:2505.15814
-
AdUE: Improving uncertainty estimation head for LoRA adapters in LLMs 21 May 2025 · 0 repositories · arXiv:2505.15443
-
An Efficient Private GPT Never Autoregressively Decodes 21 May 2025 · 0 repositories · arXiv:2505.15252
-
An Exploratory Approach Towards Investigating and Explaining Vision Transformer and Transfer Learning for Brain Disease Detection 21 May 2025 · 0 repositories · arXiv:2505.16039
-
Beyond Node Attention: Multi-Scale Harmonic Encoding for Feature-Wise Graph Message Passing 21 May 2025 · 0 repositories · arXiv:2505.15015
-
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems 21 May 2025 · 0 repositories · arXiv:2505.15216
-
BR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law 21 May 2025 · 0 repositories · arXiv:2505.15916
-
CEBSNet: Change-Excited and Background-Suppressed Network with Temporal Dependency Modeling for Bitemporal Change Detection 21 May 2025 · 0 repositories · arXiv:2505.15322
-
Collaborative Problem-Solving in an Optimization Game 21 May 2025 · 1 repository · arXiv:2505.15490
-
Convolutional Long Short-Term Memory Neural Networks Based Numerical Simulation of Flow Field 21 May 2025 · 0 repositories · arXiv:2505.15533
-
Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging 21 May 2025 · 0 repositories · arXiv:2505.15875
-
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 21 May 2025 · 0 repositories · arXiv:2505.15045
-
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data 21 May 2025 · 0 repositories · arXiv:2505.15074
-
dKV-Cache: The Cache for Diffusion Language Models 21 May 2025 · 2 repositories · arXiv:2505.15781
-
Domain Adaptive Skin Lesion Classification via Conformal Ensemble of Vision Transformers 21 May 2025 · 0 repositories · arXiv:2505.15997
-
Exploring the Innovation Opportunities for Pre-trained Models 21 May 2025 · 0 repositories · arXiv:2505.15790
-
Filtering Learning Histories Enhances In-Context Reinforcement Learning 21 May 2025 · 0 repositories · arXiv:2505.15143
-
Fourier-Invertible Neural Encoder (FINE) for Homogeneous Flows 21 May 2025 · 0 repositories · arXiv:2505.15329
-
Gated Integration of Low-Rank Adaptation for Continual Learning of Language Models 21 May 2025 · 1 repository · arXiv:2505.15424Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Guidelines for the Quality Assessment of Energy-Aware NAS Benchmarks 21 May 2025 · 0 repositories · arXiv:2505.15631
-
Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization 21 May 2025 · 0 repositories · arXiv:2505.15291