Methods › General › Attention Mechanisms › Attention › Papers, page 11
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 11 of 316: papers 1,001 to 1,100 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation 18 May 2025 · 0 repositories · arXiv:2505.12486
-
K-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks 18 May 2025 · 0 repositories · arXiv:2505.12268
-
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection 18 May 2025 · 1 repository · arXiv:2505.14714
-
Mutual Evidential Deep Learning for Medical Image Segmentation 18 May 2025 · 0 repositories · arXiv:2505.12418
-
Near-Optimal Sample Complexities of Divergence-based S-rectangular Distributionally Robust Reinforcement Learning 18 May 2025 · 0 repositories · arXiv:2505.12202
-
PoisonArena: Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation 18 May 2025 · 1 repository · arXiv:2505.12574
-
RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines 18 May 2025 · 0 repositories · arXiv:2505.13538
-
SchoenbAt: Rethinking Attention with Polynomial basis 18 May 2025 · 1 repository · arXiv:2505.12252
-
SenseFlow: A Physics-Informed and Self-Ensembling Iterative Framework for Power Flow Estimation 18 May 2025 · 0 repositories · arXiv:2505.12302
-
SMFusion: Semantic-Preserving Fusion of Multimodal Medical Images for Enhanced Clinical Diagnosis 18 May 2025 · 0 repositories · arXiv:2505.12251
-
SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks 18 May 2025 · 0 repositories · arXiv:2505.12292
-
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference 18 May 2025 · 0 repositories · arXiv:2505.12359
-
Stereographic Multi-Try Metropolis Algorithms for Heavy-tailed Sampling 18 May 2025 · 0 repositories · arXiv:2505.12487
-
Temporal-Spectral-Spatial Unified Remote Sensing Dense Prediction 18 May 2025 · 1 repository · arXiv:2505.12280
-
Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled 18 May 2025 · 0 repositories · arXiv:2505.12196
-
Video-GPT via Next Clip Diffusion 18 May 2025 · 1 repository · arXiv:2505.12489
-
VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning 18 May 2025 · 0 repositories · arXiv:2505.12332
-
Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling 17 May 2025 · 0 repositories · arXiv:2505.12048
-
AdaptMol: Adaptive Fusion from Sequence String to Topological Structure for Few-shot Drug Discovery 17 May 2025 · 0 repositories · arXiv:2505.11878
-
Black-box Adversaries from Latent Space: Unnoticeable Attacks on Human Pose and Shape Estimation 17 May 2025 · 0 repositories · arXiv:2505.12009
-
Chain-of-Model Learning for Language Model 17 May 2025 · 0 repositories · arXiv:2505.11820
-
CL-CaGAN: Capsule differential adversarial continuous learning for cross-domain hyperspectral anomaly detection 17 May 2025 · 0 repositories · arXiv:2505.11793
-
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance 17 May 2025 · 1 repository · arXiv:2505.14708
-
ELITE: Embedding-Less retrieval with Iterative Text Exploration 17 May 2025 · 1 repository · arXiv:2505.11908
-
Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning 17 May 2025 · 0 repositories · arXiv:2505.11922
-
Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform 17 May 2025 · 0 repositories · arXiv:2505.11892
-
FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge 17 May 2025 · 1 repository · arXiv:2505.14709
-
FL-PLAS: Federated Learning with Partial Layer Aggregation for Backdoor Defense Against High-Ratio Malicious Clients 17 May 2025 · 1 repository · arXiv:2505.12019
-
GeoMaNO: Geometric Mamba Neural Operator for Partial Differential Equations 17 May 2025 · 0 repositories · arXiv:2505.12020
-
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models 17 May 2025 · 0 repositories · arXiv:2505.13514
-
Learning High-Order Relationships with Hypergraph Attention-based Spatio-Temporal Aggregation for Brain Disease Analysis 17 May 2025 · 0 repositories · arXiv:2505.12068
-
Learning to Dissipate Energy in Oscillatory State-Space Models 17 May 2025 · 1 repository · arXiv:2505.12171
-
Let's have a chat with the EU AI Act 17 May 2025 · 0 repositories · arXiv:2505.11946
-
Lightweight Spatio-Temporal Attention Network with Graph Embedding and Rotational Position Encoding for Traffic Forecasting 17 May 2025 · 0 repositories · arXiv:2505.12136
-
LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades 17 May 2025 · 0 repositories · arXiv:2505.13515
-
MedVKAN: Efficient Feature Extraction with Mamba and KAN for Medical Image Segmentation 17 May 2025 · 1 repository · arXiv:2505.11797
-
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models 17 May 2025 · 1 repository · arXiv:2505.17061
-
Neuro-Symbolic Query Compiler 17 May 2025 · 1 repository · arXiv:2505.11932
-
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations 17 May 2025 · 0 repositories · arXiv:2505.11992
-
Telco-oRAG: Optimizing Retrieval-augmented Generation for Telecom Queries via Hybrid Retrieval and Neural Routing 17 May 2025 · 0 repositories · arXiv:2505.11856
-
The Logical Expressiveness of Temporal GNNs via Two-Dimensional Product Logics 17 May 2025 · 0 repositories · arXiv:2505.11930
-
Towards Comprehensive Argument Analysis in Education: Dataset, Tasks, and Method 17 May 2025 · 0 repositories · arXiv:2505.12028
-
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation 17 May 2025 · 0 repositories · arXiv:2505.11995
-
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation 17 May 2025 · 1 repository · arXiv:2505.11849
-
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement 17 May 2025 · 1 repository · arXiv:2505.12060Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Attention-Based Reward Shaping for Sparse and Delayed Rewards 16 May 2025 · 1 repository · arXiv:2505.10802
-
A High-Performance Thermal Infrared Object Detection Framework with Centralized Regulation 16 May 2025 · 0 repositories · arXiv:2505.10825
-
RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objects 16 May 2025 · 0 repositories · arXiv:2505.10841
-
Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks? 16 May 2025 · 0 repositories · arXiv:2505.10862
-
CTP: A hybrid CNN-Transformer-PINN model for ocean front forecasting 16 May 2025 · 0 repositories · arXiv:2505.10894
-
Automated Identification of Logical Errors in Programs: Advancing Scalable Analysis of Student Misconceptions 16 May 2025 · 0 repositories · arXiv:2505.10913
-
Connecting the Dots: A Chain-of-Collaboration Prompting Framework for LLM Agents 16 May 2025 · 0 repositories · arXiv:2505.10936
-
SubGCache: Accelerating Graph-based RAG with Subgraph-level KV Cache 16 May 2025 · 0 repositories · arXiv:2505.10951
-
Relational Graph Transformer 16 May 2025 · 1 repository · arXiv:2505.10960Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
RAGSynth: Synthetic Data for Robust and Faithful RAG Component Optimization 16 May 2025 · 1 repository · arXiv:2505.10989
-
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning 16 May 2025 · 1 repository · arXiv:2505.11004
-
Rethinking the Mean Teacher Strategy from the Perspective of Self-paced Learning 16 May 2025 · 0 repositories · arXiv:2505.11018
-
Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling Scenarios 16 May 2025 · 0 repositories · arXiv:2505.11035
-
Efficient Attention via Pre-Scoring: Prioritizing Informative Keys in Transformers 16 May 2025 · 1 repository · arXiv:2505.11040
-
ShiQ: Bringing back Bellman to LLMs 16 May 2025 · 0 repositories · arXiv:2505.11081
-
Fault Diagnosis across Heterogeneous Domains via Self-Adaptive Temporal-Spatial Attention and Sample Generation 16 May 2025 · 1 repository · arXiv:2505.11083
-
Redundancy-Aware Pretraining of Vision-Language Foundation Models in Remote Sensing 16 May 2025 · 0 repositories · arXiv:2505.11121
-
GraphOracle: A Foundation Model for Knowledge Graph Reasoning 16 May 2025 · 0 repositories · arXiv:2505.11125
-
STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible Benchmarking 16 May 2025 · 1 repository · arXiv:2505.11151
-
Attention on the Sphere 16 May 2025 · 1 repository · arXiv:2505.11157
-
Maximizing Asynchronicity in Event-based Neural Networks 16 May 2025 · 0 repositories · arXiv:2505.11165
-
CheX-DS: Improving Chest X-ray Image Classification with Ensemble Learning Based on DenseNet and Swin Transformer 16 May 2025 · 0 repositories · arXiv:2505.11168
-
mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs 16 May 2025 · 1 repository · arXiv:2505.11180
-
DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling 16 May 2025 · 1 repository · arXiv:2505.11196Syntology official (archive's flag): 4 ran · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
NoPE: The Counting Power of Transformers with No Positional Encodings 16 May 2025 · 0 repositories · arXiv:2505.11199
-
GLOVA: Global and Local Variation-Aware Analog Circuit Design with Risk-Sensitive Reinforcement Learning 16 May 2025 · 0 repositories · arXiv:2505.11208
-
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction 16 May 2025 · 0 repositories · arXiv:2505.11254
-
Fractal Graph Contrastive Learning 16 May 2025 · 0 repositories · arXiv:2505.11356
-
LGBQPC: Local Granular-Ball Quality Peaks Clustering 16 May 2025 · 0 repositories · arXiv:2505.11359
-
Towards Cultural Bridge by Bahnaric-Vietnamese Translation Using Transfer Learning of Sequence-To-Sequence Pre-training Language Model 16 May 2025 · 0 repositories · arXiv:2505.11421
-
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs 16 May 2025 · 0 repositories · arXiv:2505.11423
-
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production 16 May 2025 · 0 repositories · arXiv:2505.11432
-
A Classical View on Benign Overfitting: The Role of Sample Size 16 May 2025 · 0 repositories · arXiv:2505.11621
-
ACSE-Eval: Can LLMs threat model real-world cloud infrastructure? 16 May 2025 · 1 repository · arXiv:2505.11565
-
AI-Driven Digital Transformation and Firm Performance in Chinese Industrial Enterprises: Mediating Role of Green Digital Innovation and Moderating Effects of Human-AI Collaboration 16 May 2025 · 0 repositories · arXiv:2505.11558
-
Can an Easy-to-Hard Curriculum Make Reasoning Emerge in Small Language Models? Evidence from a Four-Stage Curriculum on GPT-2 16 May 2025 · 0 repositories · arXiv:2505.11643
-
EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generation 16 May 2025 · 0 repositories · arXiv:2505.13506
-
Enhancing Mathematics Learning for Hard-of-Hearing Students Through Real-Time Palestinian Sign Language Recognition: A New Dataset 16 May 2025 · 0 repositories · arXiv:2505.17055
-
Finetune-RAG: Fine-Tuning Language Models to Resist Hallucination in Retrieval-Augmented Generation 16 May 2025 · 1 repository · arXiv:2505.10792
-
Flash Invariant Point Attention 16 May 2025 · 1 repository · arXiv:2505.11580
-
Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors 16 May 2025 · 1 repository · arXiv:2505.11612
-
Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents 16 May 2025 · 0 repositories · arXiv:2505.10961
-
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation 16 May 2025 · 1 repository · arXiv:2505.11754
-
Mathematical models for the EP2 and EP4 signaling pathways and their crosstalk 16 May 2025 · 0 repositories · arXiv:2505.11712
-
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency 16 May 2025 · 0 repositories · arXiv:2505.13499
-
Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks 16 May 2025 · 0 repositories · arXiv:2505.10909
-
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training 16 May 2025 · 1 repository · arXiv:2505.11594Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering 16 May 2025 · 0 repositories · arXiv:2505.11626
-
Transforming Decoder-Only Transformers for Accurate WiFi-Telemetry Based Indoor Localization 16 May 2025 · 0 repositories · arXiv:2505.15835
-
Vaiage: A Multi-Agent Solution to Personalized Travel Planning 16 May 2025 · 0 repositories · arXiv:2505.10922
-
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training 16 May 2025 · 0 repositories · arXiv:2505.11739
-
ARFC-WAHNet: Adaptive Receptive Field Convolution and Wavelet-Attentive Hierarchical Network for Infrared Small Target Detection 15 May 2025 · 1 repository · arXiv:2505.10595
-
SRMamba: Mamba for Super-Resolution of LiDAR Point Clouds 15 May 2025 · 0 repositories · arXiv:2505.10601
-
Continuity and Isolation Lead to Doubts or Dilemmas in Large Language Models 15 May 2025 · 0 repositories · arXiv:2505.10606
-
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly 15 May 2025 · 1 repository · arXiv:2505.10610Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)