Methods › General › Attention Mechanisms › Attention › Papers, page 45
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 45 of 316: papers 4,401 to 4,500 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Novel Convolutional-Free Method for 3D Medical Imaging Segmentation 8 Feb 2025 · 0 repositories · arXiv:2502.05396
-
APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding 8 Feb 2025 · 1 repository · arXiv:2502.05431Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Bridging Traffic State and Trajectory for Dynamic Road Network and Trajectory Representation Learning 8 Feb 2025 · 1 repository · arXiv:2502.06870
-
Diffusion Model for Interest Refinement in Multi-Interest Recommendation 8 Feb 2025 · 0 repositories · arXiv:2502.05561
-
Event Stream-based Visual Object Tracking: HDETrack V2 and A High-Definition Benchmark 8 Feb 2025 · 1 repository · arXiv:2502.05574
-
Flow-based Conformal Prediction for Multi-dimensional Time Series 8 Feb 2025 · 0 repositories · arXiv:2502.05709
-
Flowing Through Layers: A Continuous Dynamical Systems Perspective on Transformers 8 Feb 2025 · 0 repositories · arXiv:2502.05656
-
Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests 8 Feb 2025 · 0 repositories · arXiv:2502.06867
-
Graph Neural Network Enabled Pinching Antennas 8 Feb 2025 · 0 repositories · arXiv:2502.05447
-
GWRF: A Generalizable Wireless Radiance Field for Wireless Signal Propagation Modeling 8 Feb 2025 · 0 repositories · arXiv:2502.05708
-
Hierarchical Lexical Manifold Projection in Large Language Models: A Novel Mechanism for Multi-Scale Semantic Representation 8 Feb 2025 · 0 repositories · arXiv:2502.05395
-
Knowledge Graph-Guided Retrieval Augmented Generation 8 Feb 2025 · 1 repository · arXiv:2502.06864
-
Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding 8 Feb 2025 · 1 repository · arXiv:2502.05609
-
MoFM: A Large-Scale Human Motion Foundation Model 8 Feb 2025 · 0 repositories · arXiv:2502.05432
-
Multi-scale Masked Autoencoder for Electrocardiogram Anomaly Detection 8 Feb 2025 · 0 repositories · arXiv:2502.05494
-
NOMANet: A Graph Neural Network Enabled Power Allocation Scheme for NOMA 8 Feb 2025 · 0 repositories · arXiv:2502.05592
-
On the Effectiveness of Large Language Models in Automating Categorization of Scientific Texts 8 Feb 2025 · 0 repositories · arXiv:2502.15745
-
Probabilistic Foundations for Metacognition via Hybrid-AI 8 Feb 2025 · 0 repositories · arXiv:2502.05398
-
TabICL: A Tabular Foundation Model for In-Context Learning on Large Data 8 Feb 2025 · 0 repositories · arXiv:2502.05564
-
The Odyssey of the Fittest: Can Agents Survive and Still Be Good? 8 Feb 2025 · 1 repository · arXiv:2502.05442
-
Topological derivative approach for deep neural network architecture adaptation 8 Feb 2025 · 0 repositories · arXiv:2502.06885
-
Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey 8 Feb 2025 · 1 repository · arXiv:2502.06872
-
A Deep Learning Framework Integrating CNN and BiLSTM for Financial Systemic Risk Analysis and Prediction 7 Feb 2025 · 0 repositories · arXiv:2502.06847
-
Can Large Language Models Understand Intermediate Representations? 7 Feb 2025 · 0 repositories · arXiv:2502.06854
-
Cross-Encoder Rediscovers a Semantic Variant of BM25 7 Feb 2025 · 0 repositories · arXiv:2502.04645
-
Detection of LLM-Generated Java Code Using Discretized Nested Bigrams 7 Feb 2025 · 0 repositories · arXiv:2502.15740
-
Discounting under inequality and lobbyists disagreement 7 Feb 2025 · 0 repositories · arXiv:2502.05342
-
DMPA: Model Poisoning Attacks on Decentralized Federated Learning for Model Differences 7 Feb 2025 · 0 repositories · arXiv:2502.04771
-
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification 7 Feb 2025 · 0 repositories · arXiv:2502.06852
-
Efficient Knowledge Feeding to Language Models: A Novel Integrated Encoder-Decoder Architecture 7 Feb 2025 · 0 repositories · arXiv:2502.05233
-
Enhancing Pre-Trained Decision Transformers with Prompt-Tuning Bandits 7 Feb 2025 · 0 repositories · arXiv:2502.04979
-
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics? 7 Feb 2025 · 0 repositories · arXiv:2502.04718
-
HetSSNet: Spatial-Spectral Heterogeneous Graph Learning Network for Panchromatic and Multispectral Images Fusion 7 Feb 2025 · 0 repositories · arXiv:2502.04623
-
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation 7 Feb 2025 · 0 repositories · arXiv:2502.04847
-
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval 7 Feb 2025 · 0 repositories · arXiv:2502.05164
-
Is attention all you need to solve the correlated electron problem? 7 Feb 2025 · 0 repositories · arXiv:2502.05383
-
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance 7 Feb 2025 · 0 repositories · arXiv:2502.05236
-
Learning the Language of NVMe Streams for Ransomware Detection 7 Feb 2025 · 0 repositories · arXiv:2502.05011
-
Limited attention and models of choice: A behavioral equivalence 7 Feb 2025 · 0 repositories · arXiv:2502.14879
-
LP-DETR: Layer-wise Progressive Relations for Object Detection 7 Feb 2025 · 0 repositories · arXiv:2502.05147
-
MedMimic: Physician-Inspired Multimodal Fusion for Early Diagnosis of Fever of Unknown Origin 7 Feb 2025 · 0 repositories · arXiv:2502.04794
-
NoLiMa: Long-Context Evaluation Beyond Literal Matching 7 Feb 2025 · 1 repository · arXiv:2502.05167
-
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers 7 Feb 2025 · 0 repositories · arXiv:2502.05076
-
Probing Internal Representations of Multi-Word Verbs in Large Language Models 7 Feb 2025 · 0 repositories · arXiv:2502.04789
-
Representational Alignment with Chemical Induced Fit for Molecular Relational Learning 7 Feb 2025 · 0 repositories · arXiv:2502.07027
-
SelaFD:Seamless Adaptation of Vision Transformer Fine-tuning for Radar-based Human Activity 7 Feb 2025 · 1 repository · arXiv:2502.04740
-
Swin-MSTP: Swin transformer with multi-scale temporal perception for continuous sign language recognition 7 Feb 2025 · 1 repository
-
The "negative end" of change in grammar: terminology, concepts and causes 7 Feb 2025 · 0 repositories · arXiv:2502.04729
-
Wavelet-Assisted Multi-Frequency Attention Network for Pansharpening 7 Feb 2025 · 1 repository · arXiv:2502.04903
-
A Classification System Approach in Predicting Chinese Censorship 6 Feb 2025 · 0 repositories · arXiv:2502.04234
-
A Decoding Algorithm for Length-Control Summarization Based on Directed Acyclic Transformers 6 Feb 2025 · 1 repository · arXiv:2502.04535
-
A Retrospective Systematic Study on Hierarchical Sparse Query Transformer-assisted Ultrasound Screening for Early Hepatocellular Carcinoma 6 Feb 2025 · 1 repository · arXiv:2502.03772
-
A Self-supervised Multimodal Deep Learning Approach to Differentiate Post-radiotherapy Progression from Pseudoprogression in Glioblastoma 6 Feb 2025 · 0 repositories · arXiv:2502.03999
-
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference 6 Feb 2025 · 1 repository · arXiv:2502.04077
-
Beyond the Final Layer: Hierarchical Query Fusion Transformer with Agent-Interpolation Initialization for 3D Instance Segmentation 6 Feb 2025 · 0 repositories · arXiv:2502.04139
-
Boosting Knowledge Graph-based Recommendations through Confidence-Aware Augmentation with Large Language Models 6 Feb 2025 · 0 repositories · arXiv:2502.03715
-
Building A Unified AI-centric Language System: analysis, framework and future work 6 Feb 2025 · 0 repositories · arXiv:2502.04488
-
CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction 6 Feb 2025 · 0 repositories · arXiv:2502.06836
-
CleanSurvival: Automated data preprocessing for time-to-event models using reinforcement learning 6 Feb 2025 · 0 repositories · arXiv:2502.03946
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features 6 Feb 2025 · 1 repository · arXiv:2502.04320Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
DEALing with Image Reconstruction: Deep Attentive Least Squares 6 Feb 2025 · 0 repositories · arXiv:2502.04079
-
Decoding Human Attentive States from Spatial-temporal EEG Patches Using Transformers 6 Feb 2025 · 1 repository · arXiv:2502.03736
-
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation 6 Feb 2025 · 0 repositories · arXiv:2502.03930
-
Expanding Training Data for Endoscopic Phenotyping of Eosinophilic Esophagitis 6 Feb 2025 · 0 repositories · arXiv:2502.04199
-
Experiments with Large Language Models on Retrieval-Augmented Generation for Closed-Source Simulation Software 6 Feb 2025 · 0 repositories · arXiv:2502.03916
-
Fast Video Generation with Sliding Tile Attention 6 Feb 2025 · 1 repository · arXiv:2502.04507
-
Graph machine learning for flight delay prediction due to holding manouver 6 Feb 2025 · 1 repository · arXiv:2502.04233
-
How vulnerable is my policy? Adversarial attacks on modern behavior cloning policies 6 Feb 2025 · 0 repositories · arXiv:2502.03698
-
ICGNN: Graph Neural Network Enabled Scalable Beamforming for MISO Interference Channels 6 Feb 2025 · 0 repositories · arXiv:2502.03936
-
Identify Critical KV Cache in LLM Inference from an Output Perturbation Perspective 6 Feb 2025 · 2 repositories · arXiv:2502.03805Syntology official (archive's flag): 1 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 11 harvested samples)
-
ImprovNet -- Generating Controllable Musical Improvisations with Iterative Corruption Refinement 6 Feb 2025 · 1 repository · arXiv:2502.04522
-
KVTuner: Sensitivity-Aware Layer-wise Mixed Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference 6 Feb 2025 · 1 repository · arXiv:2502.04420
-
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis 6 Feb 2025 · 1 repository · arXiv:2502.04128
-
LLMs to Support a Domain Specific Knowledge Assistant 6 Feb 2025 · 0 repositories · arXiv:2502.04095
-
MD-BERT: Action Recognition in Dark Videos via Dynamic Multi-Stream Fusion and Temporal Modeling 6 Feb 2025 · 1 repository · arXiv:2502.03724
-
MedGNN: Towards Multi-resolution Spatiotemporal Graph Learning for Medical Time Series Classification 6 Feb 2025 · 1 repository · arXiv:2502.04515
-
MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare Copilot 6 Feb 2025 · 1 repository · arXiv:2502.04413
-
MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation 6 Feb 2025 · 1 repository · arXiv:2502.04176
-
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation 6 Feb 2025 · 1 repository · arXiv:2502.04537
-
No Images, No Problem: Retaining Knowledge in Continual VQA with Questions-Only Memory 6 Feb 2025 · 1 repository · arXiv:2502.04469
-
On the importance of structural identifiability for machine learning with partially observed dynamical systems 6 Feb 2025 · 1 repository · arXiv:2502.04131
-
Optimized Unet with Attention Mechanism for Multi-Scale Semantic Segmentation 6 Feb 2025 · 0 repositories · arXiv:2502.03813
-
Position: Untrained Machine Learning for Anomaly Detection 6 Feb 2025 · 0 repositories · arXiv:2502.03876
-
Private Federated Learning In Real World Application -- A Case Study 6 Feb 2025 · 0 repositories · arXiv:2502.04565
-
PsyPlay: Personality-Infused Role-Playing Conversational Agents 6 Feb 2025 · 0 repositories · arXiv:2502.03821
-
Semantic Feature Division Multiple Access for Digital Semantic Broadcast Channels 6 Feb 2025 · 0 repositories · arXiv:2502.03949
-
SMI: An Information-Theoretic Metric for Predicting Model Knowledge Solely from Pre-Training Signals 6 Feb 2025 · 1 repository · arXiv:2502.04066
-
SWIPTNet: A Unified Deep Learning Framework for SWIPT based on GNN and Transfer Learning 6 Feb 2025 · 0 repositories · arXiv:2502.03928
-
UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation 6 Feb 2025 · 0 repositories · arXiv:2502.04393
-
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation 6 Feb 2025 · 0 repositories · arXiv:2502.06843
-
Zero-shot Meta-learning for Tabular Prediction Tasks with Adversarially Pre-trained Transformer 6 Feb 2025 · 0 repositories · arXiv:2502.04573
-
A Contemporary Survey of Large Language Model Assisted Program Analysis 5 Feb 2025 · 0 repositories · arXiv:2502.18474
-
A Schema-Guided Reason-while-Retrieve framework for Reasoning on Scene Graphs with Large-Language-Models (LLMs) 5 Feb 2025 · 0 repositories · arXiv:2502.03450
-
Accessible and Portable LLM Inference by Compiling Computational Graphs into SQL 5 Feb 2025 · 0 repositories · arXiv:2502.02818
-
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training 5 Feb 2025 · 0 repositories · arXiv:2502.03460
-
Analyzing limits for in-context learning 5 Feb 2025 · 0 repositories · arXiv:2502.03503
-
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation 5 Feb 2025 · 0 repositories · arXiv:2502.15734
-
CTR-Driven Advertising Image Generation with Multimodal Large Language Models 5 Feb 2025 · 1 repository · arXiv:2502.06823
-
DC-VSR: Spatially and Temporally Consistent Video Super-Resolution with Video Diffusion Prior 5 Feb 2025 · 0 repositories · arXiv:2502.03502
-
DynVFX: Augmenting Real Videos with Dynamic Content 5 Feb 2025 · 0 repositories · arXiv:2502.03621