Methods › General › Attention Mechanisms › Attention › Papers, page 40
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 40 of 316: papers 3,901 to 4,000 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation 21 Feb 2025 · 0 repositories · arXiv:2502.15203
-
Generalization Guarantees for Representation Learning via Data-Dependent Gaussian Mixture Priors 21 Feb 2025 · 1 repository · arXiv:2502.15540Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
GNN-Coder: Boosting Semantic Code Retrieval with Combined GNNs and Transformer 21 Feb 2025 · 0 repositories · arXiv:2502.15202
-
Graph Attention Convolutional U-NET: A Semantic Segmentation Model for Identifying Flooded Areas 21 Feb 2025 · 0 repositories · arXiv:2502.15907
-
Graph-Based Deep Learning on Stereo EEG for Predicting Seizure Freedom in Epilepsy Patients 21 Feb 2025 · 0 repositories · arXiv:2502.15198
-
Interleaved Block-based Learned Image Compression with Feature Enhancement and Quantization Error Compensation 21 Feb 2025 · 0 repositories · arXiv:2502.15188
-
Jeffrey's update rule as a minimizer of Kullback-Leibler divergence 21 Feb 2025 · 0 repositories · arXiv:2502.15504
-
Learning with Limited Shared Information in Multi-agent Multi-armed Bandit 21 Feb 2025 · 0 repositories · arXiv:2502.15338
-
LightThinker: Thinking Step-by-Step Compression 21 Feb 2025 · 0 repositories · arXiv:2502.15589
-
Lightweight yet Efficient: An External Attentive Graph Convolutional Network with Positional Prompts for Sequential Recommendation 21 Feb 2025 · 1 repository · arXiv:2502.15331
-
LUMINA-Net: Low-light Upgrade through Multi-stage Illumination and Noise Adaptation Network for Image Enhancement 21 Feb 2025 · 0 repositories · arXiv:2502.15186
-
M2LADS Demo: A System for Generating Multimodal Learning Analytics Dashboards 21 Feb 2025 · 0 repositories · arXiv:2502.15363
-
Mantis: Lightweight Calibrated Foundation Model for User-Friendly Time Series Classification 21 Feb 2025 · 1 repository · arXiv:2502.15637Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models 21 Feb 2025 · 0 repositories · arXiv:2502.15950
-
Retrieval-Augmented Speech Recognition Approach for Domain Challenges 21 Feb 2025 · 0 repositories · arXiv:2502.15264
-
Robust Bias Detection in MLMs and its Application to Human Trait Ratings 21 Feb 2025 · 1 repository · arXiv:2502.15600
-
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference 21 Feb 2025 · 0 repositories · arXiv:2502.15294
-
SentiFormer: Metadata Enhanced Transformer for Image Sentiment Analysis 21 Feb 2025 · 1 repository · arXiv:2502.15322
-
Single-pass Detection of Jailbreaking Input in Large Language Models 21 Feb 2025 · 0 repositories · arXiv:2502.15435
-
Soybean pod and seed counting in both outdoor fields and indoor laboratories using unions of deep neural networks 21 Feb 2025 · 0 repositories · arXiv:2502.15286
-
Sparks of cognitive flexibility: self-guided context inference for flexible stimulus-response mapping by attentional routing 21 Feb 2025 · 0 repositories · arXiv:2502.15634
-
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention 21 Feb 2025 · 0 repositories · arXiv:2502.15304
-
Tokenization is Sensitive to Language Variation 21 Feb 2025 · 0 repositories · arXiv:2502.15343
-
TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba 21 Feb 2025 · 0 repositories · arXiv:2502.15130
-
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice 21 Feb 2025 · 1 repository · arXiv:2502.18504
-
Utilizing Sequential Information of General Lab-test Results and Diagnoses History for Differential Diagnosis of Dementia 21 Feb 2025 · 0 repositories · arXiv:2502.15317
-
A Socratic RAG Approach to Connect Natural Language Queries on Research Topics with Knowledge Organization Systems 20 Feb 2025 · 0 repositories · arXiv:2502.15005
-
Adaptive Convolution for CNN-based Speech Enhancement Models 20 Feb 2025 · 1 repository · arXiv:2502.14224
-
Argument-Based Comparative Question Answering Evaluation Benchmark 20 Feb 2025 · 0 repositories · arXiv:2502.14476
-
Asymmetric Co-Training for Source-Free Few-Shot Domain Adaptation 20 Feb 2025 · 1 repository · arXiv:2502.14214
-
Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place Recognition 20 Feb 2025 · 1 repository · arXiv:2502.14195
-
Cardiac Evidence Backtracking for Eating Behavior Monitoring using Collocative Electrocardiogram Imagining 20 Feb 2025 · 0 repositories · arXiv:2502.14430
-
DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model 20 Feb 2025 · 0 repositories · arXiv:2502.15832
-
Designing Parameter and Compute Efficient Diffusion Transformers using Distillation 20 Feb 2025 · 0 repositories · arXiv:2502.14226
-
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions 20 Feb 2025 · 0 repositories · arXiv:2502.14202
-
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information 20 Feb 2025 · 1 repository · arXiv:2502.14258
-
Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective 20 Feb 2025 · 1 repository · arXiv:2502.14340Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Effects of Prompt Length on Domain-specific Tasks for Large Language Models 20 Feb 2025 · 0 repositories · arXiv:2502.14255
-
Entropy-UID: A Method for Optimizing Information Density 20 Feb 2025 · 0 repositories · arXiv:2502.14366
-
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity 20 Feb 2025 · 1 repository · arXiv:2502.14620
-
FIND: Fine-grained Information Density Guided Adaptive Retrieval-Augmented Generation for Disease Diagnosis 20 Feb 2025 · 0 repositories · arXiv:2502.14614
-
Forecasting Local Ionospheric Parameters Using Transformers 20 Feb 2025 · 1 repository · arXiv:2502.15093
-
From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT 20 Feb 2025 · 0 repositories · arXiv:2502.14714
-
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models 20 Feb 2025 · 1 repository · arXiv:2502.14802Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Model Inversion Attack against Federated Unlearning 20 Feb 2025 · 0 repositories · arXiv:2502.14558
-
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning 20 Feb 2025 · 0 repositories · arXiv:2502.14356
-
H3DE-Net: Efficient and Accurate 3D Landmark Detection in Medical Imaging 20 Feb 2025 · 1 repository · arXiv:2502.14221
-
Hallucination Detection in Large Language Models with Metamorphic Relations 20 Feb 2025 · 0 repositories · arXiv:2502.15844
-
Hardware-Friendly Static Quantization Method for Video Diffusion Transformers 20 Feb 2025 · 0 repositories · arXiv:2502.15077
-
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation 20 Feb 2025 · 1 repository · arXiv:2502.14642
-
Is Relevance Propagated from Retriever to Generator in RAG? 20 Feb 2025 · 0 repositories · arXiv:2502.15025
-
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding 20 Feb 2025 · 0 repositories · arXiv:2502.14949
-
LIFT: Improving Long Context Understanding of Large Language Models through Long Input Fine-Tuning 20 Feb 2025 · 0 repositories · arXiv:2502.14644
-
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention 20 Feb 2025 · 2 repositories · arXiv:2502.14866Syntology official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Mechanistic Understanding of Language Models in Syntactic Code Completion 20 Feb 2025 · 0 repositories · arXiv:2502.18499
-
Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling 20 Feb 2025 · 1 repository · arXiv:2502.14553
-
NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis 20 Feb 2025 · 0 repositories · arXiv:2502.14178
-
On the Influence of Context Size and Model Choice in Retrieval-Augmented Generation Systems 20 Feb 2025 · 1 repository · arXiv:2502.14759
-
PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant 20 Feb 2025 · 0 repositories · arXiv:2502.14271
-
ParallelComp: Parallel Long-Context Compressor for Length Extrapolation 20 Feb 2025 · 0 repositories · arXiv:2502.14317
-
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models 20 Feb 2025 · 0 repositories · arXiv:2502.14504
-
Predicting Fetal Birthweight from High Dimensional Data using Advanced Machine Learning 20 Feb 2025 · 0 repositories · arXiv:2502.14270
-
QUAD-LLM-MLTC: Large Language Models Ensemble Learning for Healthcare Text Multi-Label Classification 20 Feb 2025 · 0 repositories · arXiv:2502.14189
-
Reinforcement Learning with Graph Attention for Routing and Wavelength Assignment with Lightpath Reuse 20 Feb 2025 · 0 repositories · arXiv:2502.14741
-
RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers 20 Feb 2025 · 0 repositories · arXiv:2502.14377
-
RendBEV: Semantic Novel View Synthesis for Self-Supervised Bird's Eye View Segmentation 20 Feb 2025 · 0 repositories · arXiv:2502.14792
-
Revealing and Mitigating Over-Attention in Knowledge Editing 20 Feb 2025 · 1 repository · arXiv:2502.14838
-
Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting 20 Feb 2025 · 0 repositories · arXiv:2502.15819
-
Textured 3D Regenerative Morphing with 3D Diffusion Prior 20 Feb 2025 · 0 repositories · arXiv:2502.14316
-
Topology-Aware Wavelet Mamba for Airway Structure Segmentation in Postoperative Recurrent Nasopharyngeal Carcinoma CT Scans 20 Feb 2025 · 0 repositories · arXiv:2502.14363
-
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs 20 Feb 2025 · 1 repository · arXiv:2502.14837Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples)
-
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression 20 Feb 2025 · 0 repositories · arXiv:2502.14477
-
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models 20 Feb 2025 · 0 repositories · arXiv:2502.14727
-
A consensus set for the aggregation of partial rankings: the case of the Optimal Set of Bucket Orders Problem 19 Feb 2025 · 0 repositories · arXiv:2502.13769
-
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference 19 Feb 2025 · 0 repositories · arXiv:2502.13542
-
Adapting Large Language Models for Time Series Modeling via a Novel Parameter-efficient Adaptation Method 19 Feb 2025 · 0 repositories · arXiv:2502.13725
-
Are Large Language Models In-Context Graph Learners? 19 Feb 2025 · 0 repositories · arXiv:2502.13562
-
Building Age Estimation: A New Multi-Modal Benchmark Dataset and Community Challenge 19 Feb 2025 · 1 repository · arXiv:2502.13818
-
Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning 19 Feb 2025 · 0 repositories · arXiv:2502.13754
-
Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets 19 Feb 2025 · 1 repository · arXiv:2502.13833
-
Conveniently Identify Coils in Inductive Power Transfer System Using Machine Learning 19 Feb 2025 · 0 repositories · arXiv:2502.13915
-
DH-RAG: A Dynamic Historical Context-Powered Retrieval-Augmented Generation Method for Multi-Turn Dialogue 19 Feb 2025 · 0 repositories · arXiv:2502.13847
-
Diffusion Model Agnostic Social Influence Maximization in Hyperbolic Space 19 Feb 2025 · 0 repositories · arXiv:2502.13571
-
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs 19 Feb 2025 · 0 repositories · arXiv:2502.13566
-
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference 19 Feb 2025 · 0 repositories · arXiv:2502.15804
-
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length 19 Feb 2025 · 0 repositories · arXiv:2502.13967
-
From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education 19 Feb 2025 · 0 repositories · arXiv:2502.13789
-
Generative Detail Enhancement for Physically Based Materials 19 Feb 2025 · 0 repositories · arXiv:2502.13994
-
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking 19 Feb 2025 · 0 repositories · arXiv:2502.13766
-
Giving AI Personalities Leads to More Human-Like Reasoning 19 Feb 2025 · 0 repositories · arXiv:2502.14155
-
HawkBench: Investigating Resilience of RAG Methods on Stratified Information-Seeking Tasks 19 Feb 2025 · 0 repositories · arXiv:2502.13465
-
Helix-mRNA: A Hybrid Foundation Model For Full Sequence mRNA Therapeutics 19 Feb 2025 · 1 repository · arXiv:2502.13785
-
Hidden Darkness in LLM-Generated Designs: Exploring Dark Patterns in Ecommerce Web Components Generated by LLMs 19 Feb 2025 · 0 repositories · arXiv:2502.13499
-
In-Place Updates of a Graph Index for Streaming Approximate Nearest Neighbor Search 19 Feb 2025 · 0 repositories · arXiv:2502.13826
-
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking 19 Feb 2025 · 0 repositories · arXiv:2502.13842
-
Integration of Agentic AI with 6G Networks for Mission-Critical Applications: Use-case and Challenges 19 Feb 2025 · 0 repositories · arXiv:2502.13476
-
Learning Novel Transformer Architecture for Time-series Forecasting 19 Feb 2025 · 0 repositories · arXiv:2502.13721
-
MambaLiteSR: Image Super-Resolution with Low-Rank Mamba using Knowledge Distillation 19 Feb 2025 · 0 repositories · arXiv:2502.14090
-
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures 19 Feb 2025 · 0 repositories · arXiv:2502.14008
-
Medical Image Classification with KAN-Integrated Transformers and Dilated Neighborhood Attention 19 Feb 2025 · 1 repository · arXiv:2502.13693