Methods › General › Attention Mechanisms › Attention › Papers, page 42
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 42 of 316: papers 4,101 to 4,200 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Precise GPS-Denied UAV Self-Positioning via Context-Enhanced Cross-View Geo-Localization 17 Feb 2025 · 0 repositories · arXiv:2502.11408
-
RAG vs. GraphRAG: A Systematic Evaluation and Key Insights 17 Feb 2025 · 0 repositories · arXiv:2502.11371
-
REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark 17 Feb 2025 · 0 repositories · arXiv:2502.12342
-
Market-Derived Financial Sentiment Analysis: Context-Aware Language Models for Crypto Forecasting 17 Feb 2025 · 1 repository · arXiv:2502.14897
-
Revisiting Robust RAG: Do We Still Need Complex Robust Training in the Era of Powerful LLMs? 17 Feb 2025 · 0 repositories · arXiv:2502.11400
-
S2TX: Cross-Attention Multi-Scale State-Space Transformer for Time Series Forecasting 17 Feb 2025 · 0 repositories · arXiv:2502.11340
-
SmartLLM: Smart Contract Auditing using Custom Generative AI 17 Feb 2025 · 0 repositories · arXiv:2502.13167
-
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More 17 Feb 2025 · 1 repository · arXiv:2502.11494Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs 17 Feb 2025 · 0 repositories · arXiv:2502.12216
-
The geometry of BERT 17 Feb 2025 · 0 repositories · arXiv:2502.12033
-
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It 17 Feb 2025 · 1 repository · arXiv:2502.11771
-
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models 17 Feb 2025 · 0 repositories · arXiv:2502.11458
-
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs 17 Feb 2025 · 1 repository · arXiv:2502.12352
-
Towards Practical First-Order Model Counting 17 Feb 2025 · 0 repositories · arXiv:2502.12278
-
VRoPE: Rotary Position Embedding for Video Large Language Models 17 Feb 2025 · 1 repository · arXiv:2502.11664Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
What's in a Query: Polarity-Aware Distribution-Based Fair Ranking 17 Feb 2025 · 0 repositories · arXiv:2502.11429
-
X-IL: Exploring the Design Space of Imitation Learning Policies 17 Feb 2025 · 1 repository · arXiv:2502.12330Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Zero Token-Driven Deep Thinking in LLMs: Unlocking the Full Potential of Existing Parameters via Cyclic Refinement 17 Feb 2025 · 0 repositories · arXiv:2502.12214
-
A Critical Review of Predominant Bias in Neural Networks 16 Feb 2025 · 0 repositories · arXiv:2502.11031
-
A recurrent vision transformer shows signatures of primate visual attention 16 Feb 2025 · 0 repositories · arXiv:2502.10955
-
AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks 16 Feb 2025 · 0 repositories · arXiv:2502.11158
-
AudioSpa: Spatializing Sound Events with Text 16 Feb 2025 · 0 repositories · arXiv:2502.11219
-
Bridging the Gap: Enabling Natural Language Queries for NoSQL Databases through Text-to-NoSQL Translation 16 Feb 2025 · 0 repositories · arXiv:2502.11201
-
CL-MFAP: A Contrastive Learning-Based Multimodal Foundation Model for Molecular Property Prediction and Antibiotic Screening 16 Feb 2025 · 1 repository · arXiv:2502.11001Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation 16 Feb 2025 · 1 repository · arXiv:2502.10940
-
DA-Mamba: Domain Adaptive Hybrid Mamba-Transformer Based One-Stage Object Detection 16 Feb 2025 · 2 repositories · arXiv:2502.11178
-
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions 16 Feb 2025 · 0 repositories · arXiv:2502.11073
-
DT4ECG: A Dual-Task Learning Framework for ECG-Based Human Identity Recognition and Human Activity Detection 16 Feb 2025 · 0 repositories · arXiv:2502.11023
-
Efficient Long-Decoding Inference with Reasoning-Aware Attention Sparsity 16 Feb 2025 · 0 repositories · arXiv:2502.11147
-
Empirical evaluation of LLMs in predicting fixes of Configuration bugs in Smart Home System 16 Feb 2025 · 0 repositories · arXiv:2502.10953
-
Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks 16 Feb 2025 · 0 repositories · arXiv:2502.11152
-
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models 16 Feb 2025 · 1 repository · arXiv:2502.11075
-
GS-GVINS: A Tightly-integrated GNSS-Visual-Inertial Navigation System Augmented by 3D Gaussian Splatting 16 Feb 2025 · 0 repositories · arXiv:2502.10975
-
Gumbel Reranking: Differentiable End-to-End Reranker Optimization 16 Feb 2025 · 0 repositories · arXiv:2502.11116
-
Improving Similar Case Retrieval Ranking Performance By Revisiting RankSVM 16 Feb 2025 · 1 repository · arXiv:2502.11131
-
Integrating Language Models for Enhanced Network State Monitoring in DRL-Based SFC Provisioning 16 Feb 2025 · 0 repositories · arXiv:2502.11298
-
Investigating Language Preference of Multilingual RAG Systems 16 Feb 2025 · 0 repositories · arXiv:2502.11175
-
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding 16 Feb 2025 · 1 repository · arXiv:2502.11168
-
Leveraging Conditional Mutual Information to Improve Large Language Model Fine-Tuning For Classification 16 Feb 2025 · 0 repositories · arXiv:2502.11258
-
Leveraging Constrained Monte Carlo Tree Search to Generate Reliable Long Chain-of-Thought for Mathematical Reasoning 16 Feb 2025 · 0 repositories · arXiv:2502.11169
-
MultiTEND: A Multilingual Benchmark for Natural Language to NoSQL Query Translation 16 Feb 2025 · 0 repositories · arXiv:2502.11022
-
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention 16 Feb 2025 · 3 repositories · arXiv:2502.11089Syntology 16 ran (of which 2 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 19 harvested samples) · 2 pointer-only (licence)
-
Performance Review on LLM for solving leetcode problems 16 Feb 2025 · 0 repositories · arXiv:2502.15770
-
QuOTE: Question-Oriented Text Embeddings 16 Feb 2025 · 0 repositories · arXiv:2502.10976
-
Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction 16 Feb 2025 · 1 repository · arXiv:2502.11084
-
RoseRAG: Robust Retrieval-augmented Generation with Small-scale LLMs via Margin-aware Preference Optimization 16 Feb 2025 · 0 repositories · arXiv:2502.10993
-
RT-DEMT: A hybrid real-time acupoint detection model combining mamba and transformer 16 Feb 2025 · 1 repository · arXiv:2502.11179
-
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information 16 Feb 2025 · 0 repositories · arXiv:2502.10950
-
Streamlining the Collaborative Chain of Models into A Single Forward Pass in Generation-Based Tasks 16 Feb 2025 · 1 repository · arXiv:2502.11083
-
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval 16 Feb 2025 · 0 repositories · arXiv:2502.11276
-
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages 16 Feb 2025 · 1 repository · arXiv:2502.11020
-
Attention Mechanism for LLM-based Agents Dynamic Diffusion under Information Asymmetry 16 Feb 2025 · 0 repositories · arXiv:2502.13160
-
Unveiling the Power of Complex-Valued Transformers in Wireless Communications 16 Feb 2025 · 0 repositories · arXiv:2502.11151
-
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs 16 Feb 2025 · 0 repositories · arXiv:2502.11228
-
1bit-Merging: Dynamic Quantized Merging for Large Language Models 15 Feb 2025 · 0 repositories · arXiv:2502.10743
-
Automatic Quality Assessment of First Trimester Crown-Rump-Length Ultrasound Images 15 Feb 2025 · 0 repositories · arXiv:2502.10908
-
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models 15 Feb 2025 · 0 repositories · arXiv:2502.10835
-
BalanceBenchmark: A Survey for Imbalanced Learning 15 Feb 2025 · 1 repository · arXiv:2502.10816
-
CiteCheck: Towards Accurate Citation Faithfulness Detection 15 Feb 2025 · 1 repository · arXiv:2502.10881
-
CLoCKDistill: Consistent Location-and-Context-aware Knowledge Distillation for DETRs 15 Feb 2025 · 0 repositories · arXiv:2502.10683
-
ControllableGPT: A Ground-Up Designed Controllable GPT for Molecule Optimization 15 Feb 2025 · 0 repositories · arXiv:2502.10631
-
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs 15 Feb 2025 · 0 repositories · arXiv:2502.10673
-
E2CB2former: Effecitve and Explainable Transformer for CB2 Receptor Ligand Activity Prediction 15 Feb 2025 · 0 repositories · arXiv:2502.12186
-
Evolving Hate Speech Online: An Adaptive Framework for Detection and Mitigation 15 Feb 2025 · 0 repositories · arXiv:2502.10921
-
CAE-Net: Generalized Deepfake Image Detection using Convolution and Attention Mechanisms with Spatial and Frequency Domain Features 15 Feb 2025 · 0 repositories · arXiv:2502.10682
-
HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model 15 Feb 2025 · 0 repositories · arXiv:2502.10807
-
Improving action segmentation via explicit similarity measurement 15 Feb 2025 · 0 repositories · arXiv:2502.10713
-
LEAPS: A discrete neural sampler via locally equivariant networks 15 Feb 2025 · 0 repositories · arXiv:2502.10843
-
Learning Identifiable Structures Helps Avoid Bias in DNN-based Supervised Causal Learning 15 Feb 2025 · 1 repository · arXiv:2502.10883
-
Learning to Explain Air Traffic Situation 15 Feb 2025 · 0 repositories · arXiv:2502.10764
-
Lost in the Passage: Passage-level In-context Learning Does Not Necessarily Need a "Passage" 15 Feb 2025 · 0 repositories · arXiv:2502.10634
-
NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for Personalized Hearing Aids 15 Feb 2025 · 0 repositories · arXiv:2502.10822
-
NitiBench: A Comprehensive Studies of LLM Frameworks Capabilities for Thai Legal Question Answering 15 Feb 2025 · 1 repository · arXiv:2502.10868
-
Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition 15 Feb 2025 · 0 repositories · arXiv:2502.10674
-
Order-agnostic Identifier for Large Language Model-based Generative Recommendation 15 Feb 2025 · 0 repositories · arXiv:2502.10833
-
ResiComp: Loss-Resilient Image Compression via Dual-Functional Masked Visual Token Modeling 15 Feb 2025 · 0 repositories · arXiv:2502.10812
-
Self-Explaining Hypergraph Neural Networks for Diagnosis Prediction 15 Feb 2025 · 0 repositories · arXiv:2502.10689
-
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers 15 Feb 2025 · 1 repository · arXiv:2502.10841
-
Spatio-temporal collaborative multiple-stream transformer network for liver lesion classification on multiple-sequence magnetic resonance imaging 15 Feb 2025 · 1 repository
-
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training 15 Feb 2025 · 1 repository · arXiv:2502.10927
-
VarGes: Improving Variation in Co-Speech 3D Gesture Generation via StyleCLIPS 15 Feb 2025 · 1 repository · arXiv:2502.10729
-
A Self-Supervised Reinforcement Learning Approach for Fine-Tuning Large Language Models Using Cross-Attention Signals 14 Feb 2025 · 0 repositories · arXiv:2502.10482
-
A synergistic CNN-transformer network with pooling attention fusion for hyperspectral image classification 14 Feb 2025 · 1 repository
-
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies 14 Feb 2025 · 0 repositories · arXiv:2502.09870
-
An Efficient Large Recommendation Model: Towards a Resource-Optimal Scaling Law 14 Feb 2025 · 0 repositories · arXiv:2502.09888
-
An Innovative Next Activity Prediction Approach Using Process Entropy and DAW-Transformer 14 Feb 2025 · 0 repositories · arXiv:2502.10573
-
ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation 14 Feb 2025 · 0 repositories · arXiv:2502.09891
-
Compress image to patches for Vision Transformer 14 Feb 2025 · 1 repository · arXiv:2502.10120
-
Do Large Language Models Reason Causally Like Us? Even Better? 14 Feb 2025 · 0 repositories · arXiv:2502.10215
-
EmbBERT-Q: Breaking Memory Barriers in Embedded NLP 14 Feb 2025 · 0 repositories · arXiv:2502.10001
-
Evaluating and Improving Graph-based Explanation Methods for Multi-Agent Coordination 14 Feb 2025 · 2 repositories · arXiv:2502.09889
-
From Deep Additive Kernel Learning to Last-Layer Bayesian Neural Networks via Induced Prior Approximation 14 Feb 2025 · 1 repository · arXiv:2502.10540
-
Generalized Attention Flow: Feature Attribution for Transformer Models via Maximum Flow 14 Feb 2025 · 0 repositories · arXiv:2502.15765
-
GraphiT: Efficient Node Classification on Text-Attributed Graphs with Prompt Optimized LLMs 14 Feb 2025 · 0 repositories · arXiv:2502.10522
-
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA 14 Feb 2025 · 0 repositories · arXiv:2502.10497
-
(How) Can Transformers Predict Pseudo-Random Numbers? 14 Feb 2025 · 0 repositories · arXiv:2502.10390Syntology 0 ran · 1 unverified (of 1 harvested sample)
-
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding 14 Feb 2025 · 0 repositories · arXiv:2502.09906
-
Janus: Collaborative Vision Transformer Under Dynamic Network Environment 14 Feb 2025 · 0 repositories · arXiv:2502.10047
-
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing 14 Feb 2025 · 0 repositories · arXiv:2502.09977
-
Large Language Diffusion Models 14 Feb 2025 · 2 repositories · arXiv:2502.09992Syntology 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)