Methods › General › Attention Mechanisms › Attention › Papers, page 87
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 87 of 316: papers 8,601 to 8,700 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Tending Towards Stability: Convergence Challenges in Small Language Models 15 Oct 2024 · 1 repository · arXiv:2410.11451
-
TEOcc: Radar-camera Multi-modal Occupancy Prediction via Temporal Enhancement 15 Oct 2024 · 1 repository · arXiv:2410.11228
-
Tokenization and Morphology in Multilingual Language Models: A Comparative Analysis of mT5 and ByT5 15 Oct 2024 · 0 repositories · arXiv:2410.11627
-
Towards Fair Graph Representation Learning in Social Networks 15 Oct 2024 · 0 repositories · arXiv:2410.11493
-
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings 15 Oct 2024 · 0 repositories · arXiv:2410.12046
-
TraM : Enhancing User Sleep Prediction with Transformer-based Multivariate Time Series Modeling and Machine Learning Ensembles 15 Oct 2024 · 1 repository · arXiv:2410.11293
-
Transformer Layer Injection: A Novel Approach for Efficient Upscaling of Large Language Models 15 Oct 2024 · 0 repositories · arXiv:2410.11654
-
UmambaTSF: A U-shaped Multi-Scale Long-Term Time Series Forecasting Method Using Mamba 15 Oct 2024 · 0 repositories · arXiv:2410.11278
-
Unveiling the Mystery of Visual Attributes of Concrete and Abstract Concepts: Variability, Nearest Neighbors, and Challenging Categories 15 Oct 2024 · 1 repository · arXiv:2410.11657
-
Visual Fixation-Based Retinal Prosthetic Simulation 15 Oct 2024 · 0 repositories · arXiv:2410.11688
-
YOLO-ELA: Efficient Local Attention Modeling for High-Performance Real-Time Insulator Defect Detection 15 Oct 2024 · 0 repositories · arXiv:2410.11727
-
4DStyleGaussian: Zero-shot 4D Style Transfer with Gaussian Splatting 14 Oct 2024 · 0 repositories · arXiv:2410.10412
-
A Consistency-Aware Spot-Guided Transformer for Versatile and Hierarchical Point Cloud Registration 14 Oct 2024 · 1 repository · arXiv:2410.10295
-
A few-shot Label Unlearning in Vertical Federated Learning 14 Oct 2024 · 0 repositories · arXiv:2410.10922
-
An Annotated Dataset of Errors in Premodern Greek and Baselines for Detecting Them 14 Oct 2024 · 1 repository · arXiv:2410.11071
-
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning 14 Oct 2024 · 0 repositories · arXiv:2410.10913
-
Beyond-RAG: Question Identification and Answer Generation in Real-Time Conversations 14 Oct 2024 · 0 repositories · arXiv:2410.10136
-
big.LITTLE Vision Transformer for Efficient Visual Recognition 14 Oct 2024 · 0 repositories · arXiv:2410.10267
-
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention 14 Oct 2024 · 0 repositories · arXiv:2410.10774
-
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities 14 Oct 2024 · 0 repositories · arXiv:2410.11079
-
Comparison of deep learning and conventional methods for disease onset prediction 14 Oct 2024 · 0 repositories · arXiv:2410.10505
-
Customize Your Visual Autoregressive Recipe with Set Autoregressive Modeling 14 Oct 2024 · 1 repository · arXiv:2410.10511
-
Dissecting embedding method: learning higher-order structures from data 14 Oct 2024 · 0 repositories · arXiv:2410.10917
-
Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers 14 Oct 2024 · 0 repositories · arXiv:2410.10515
-
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers 14 Oct 2024 · 1 repository · arXiv:2410.10665
-
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model 14 Oct 2024 · 0 repositories · arXiv:2410.10738
-
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads 14 Oct 2024 · 2 repositories · arXiv:2410.10819Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations 14 Oct 2024 · 1 repository · arXiv:2410.10315
-
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning 14 Oct 2024 · 0 repositories · arXiv:2410.10735
-
Enhancing Attributed Graph Networks with Alignment and Uniformity Constraints for Session-based Recommendation 14 Oct 2024 · 1 repository · arXiv:2410.10296
-
ET-Former: Efficient Triplane Deformable Attention for 3D Semantic Scene Completion From Monocular Camera 14 Oct 2024 · 0 repositories · arXiv:2410.11019
-
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification 14 Oct 2024 · 1 repository · arXiv:2410.10356Syntology 6 ran (of which 6 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 6 samples that ran constructed an object rather than computing a result (of 10 harvested samples)
-
FormalAlign: Automated Alignment Evaluation for Autoformalization 14 Oct 2024 · 1 repository · arXiv:2410.10135Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
FSOS-AMC: Few-Shot Open-Set Learning for Automatic Modulation Classification 14 Oct 2024 · 0 repositories · arXiv:2410.10265
-
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG 14 Oct 2024 · 0 repositories · arXiv:2410.10293
-
Gender Bias of LLM in Economics: An Existentialism Perspective 14 Oct 2024 · 0 repositories · arXiv:2410.19775
-
Generative AI and Its Impact on Personalized Intelligent Tutoring Systems 14 Oct 2024 · 0 repositories · arXiv:2410.10650
-
Graph of Records: Boosting Retrieval Augmented Generation for Long-context Summarization with Graphs 14 Oct 2024 · 1 repository · arXiv:2410.11001
-
GraphCLIP: Enhancing Transferability in Graph Foundation Models for Text-Attributed Graphs 14 Oct 2024 · 1 repository · arXiv:2410.10329Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer 14 Oct 2024 · 2 repositories · arXiv:2410.10812Syntology official (archive's flag): 23 ran · 23 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 2 honoured, 0 violated, 13 with no contract checked; 8 where Syntology's instrument failed) · 4 unverified (of 27 harvested samples) · 4 pointer-only (licence)
-
HSR-Enhanced Sparse Attention Acceleration 14 Oct 2024 · 0 repositories · arXiv:2410.10165
-
Hybrid Transformer for Early Alzheimer's Detection: Integration of Handwriting-Based 2D Images and 1D Signal Features 14 Oct 2024 · 0 repositories · arXiv:2410.10547
-
Interaction-Guided Two-Branch Image Dehazing Network 14 Oct 2024 · 1 repository · arXiv:2410.10121
-
KBLaM: Knowledge Base augmented Language Model 14 Oct 2024 · 1 repository · arXiv:2410.10450Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples)
-
KNN Transformer with Pyramid Prompts for Few-Shot Learning 14 Oct 2024 · 0 repositories · arXiv:2410.10227
-
Lambda-Skip Connections: the architectural component that prevents Rank Collapse 14 Oct 2024 · 0 repositories · arXiv:2410.10609
-
Learning Linear Attention in Polynomial Time 14 Oct 2024 · 0 repositories · arXiv:2410.10101
-
LKASeg:Remote-Sensing Image Semantic Segmentation with Large Kernel Attention and Full-Scale Skip Connections 14 Oct 2024 · 0 repositories · arXiv:2410.10433
-
LoLCATs: On Low-Rank Linearizing of Large Language Models 14 Oct 2024 · 1 repository · arXiv:2410.10254Syntology official (archive's flag): 23 ran · 23 ran (of which 9 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 10 where Syntology's instrument failed) · 9 unverified (of 32 harvested samples)
-
MagicEraser: Erasing Any Objects via Semantics-Aware Control 14 Oct 2024 · 1 repository · arXiv:2410.10207
-
On Calibration of LLM-based Guard Models for Reliable Content Moderation 14 Oct 2024 · 1 repository · arXiv:2410.10414Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
One Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks 14 Oct 2024 · 1 repository · arXiv:2410.11005Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Out-of-Bounding-Box Triggers: A Stealthy Approach to Cheat Object Detectors 14 Oct 2024 · 1 repository · arXiv:2410.10091Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Parameterize Structure with Differentiable Template for 3D Shape Generation 14 Oct 2024 · 0 repositories · arXiv:2410.10399
-
Performance Evaluation of Deep Learning and Transformer Models Using Multimodal Data for Breast Cancer Classification 14 Oct 2024 · 0 repositories · arXiv:2410.10146
-
Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese 14 Oct 2024 · 0 repositories · arXiv:2410.10991
-
PointNet with KAN versus PointNet with MLP for 3D Classification and Segmentation of Point Sets 14 Oct 2024 · 1 repository · arXiv:2410.10084
-
Pubic Symphysis-Fetal Head Segmentation Network Using BiFormer Attention Mechanism and Multipath Dilated Convolution 14 Oct 2024 · 0 repositories · arXiv:2410.10352
-
Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification 14 Oct 2024 · 1 repository · arXiv:2410.10573
-
Rethinking Legal Judgement Prediction in a Realistic Scenario in the Era of Large Language Models 14 Oct 2024 · 1 repository · arXiv:2410.10542
-
Reverse Refinement Network for Narrow Rural Road Detection in High-Resolution Satellite Imagery 14 Oct 2024 · 0 repositories · arXiv:2410.10389
-
Revisiting and Benchmarking Graph Autoencoders: A Contrastive Learning Perspective 14 Oct 2024 · 1 repository · arXiv:2410.10241
-
ROA-BEV: 2D Region-Oriented Attention for BEV-based 3D Object 14 Oct 2024 · 0 repositories · arXiv:2410.10298
-
RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates 14 Oct 2024 · 1 repository · arXiv:2410.10075
-
Saliency Guided Optimization of Diffusion Latents 14 Oct 2024 · 0 repositories · arXiv:2410.10257
-
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers 14 Oct 2024 · 2 repositories · arXiv:2410.10629
-
SLaNC: Static LayerNorm Calibration 14 Oct 2024 · 0 repositories · arXiv:2410.10553
-
STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with FeedBack 14 Oct 2024 · 0 repositories · arXiv:2410.10584
-
The Ingredients for Robotic Diffusion Transformers 14 Oct 2024 · 0 repositories · arXiv:2410.10088
-
Towards Better Multi-head Attention via Channel-wise Sample Permutation 14 Oct 2024 · 1 repository · arXiv:2410.10914Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Transforming Game Play: A Comparative Study of DCQN and DTQN Architectures in Reinforcement Learning 14 Oct 2024 · 0 repositories · arXiv:2410.10660
-
Transparent Networks for Multivariate Time Series 14 Oct 2024 · 1 repository · arXiv:2410.10535
-
V2M: Visual 2-Dimensional Mamba for Image Representation Learning 14 Oct 2024 · 1 repository · arXiv:2410.10382
-
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents 14 Oct 2024 · 1 repository · arXiv:2410.10594Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation 14 Oct 2024 · 1 repository · arXiv:2410.10995
-
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis 14 Oct 2024 · 1 repository · arXiv:2410.10986Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
When Attention Sink Emerges in Language Models: An Empirical View 14 Oct 2024 · 1 repository · arXiv:2410.10781Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification? 14 Oct 2024 · 1 repository · arXiv:2410.10476Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
3DS: Decomposed Difficulty Data Selection's Case Study on LLM Medical Domain Adaptation 13 Oct 2024 · 0 repositories · arXiv:2410.10901
-
A Comparative Study of PDF Parsing Tools Across Diverse Document Categories 13 Oct 2024 · 0 repositories · arXiv:2410.09871
-
BiDoRA: Bi-level Optimization-Based Weight-Decomposed Low-Rank Adaptation 13 Oct 2024 · 0 repositories · arXiv:2410.09758
-
Can In-context Learning Really Generalize to Out-of-distribution Tasks? 13 Oct 2024 · 0 repositories · arXiv:2410.09695
-
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code 13 Oct 2024 · 0 repositories · arXiv:2410.09997
-
DAG-aware Transformer for Causal Effect Estimation 13 Oct 2024 · 1 repository · arXiv:2410.10044
-
Data Adaptive Few-shot Multi Label Segmentation with Foundation Model 13 Oct 2024 · 0 repositories · arXiv:2410.09759
-
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces 13 Oct 2024 · 0 repositories · arXiv:2410.09918
-
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs 13 Oct 2024 · 1 repository · arXiv:2410.09775
-
EBDM: Exemplar-guided Image Translation with Brownian-bridge Diffusion Models 13 Oct 2024 · 0 repositories · arXiv:2410.09802
-
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation 13 Oct 2024 · 0 repositories · arXiv:2410.09704
-
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis 13 Oct 2024 · 0 repositories · arXiv:2410.12867
-
Evaluating Gender Bias of LLMs in Making Morality Judgements 13 Oct 2024 · 0 repositories · arXiv:2410.09992
-
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics 13 Oct 2024 · 1 repository · arXiv:2410.09988
-
HASN: Hybrid Attention Separable Network for Efficient Image Super-resolution 13 Oct 2024 · 1 repository · arXiv:2410.09844
-
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG 13 Oct 2024 · 0 repositories · arXiv:2410.09699
-
InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling 13 Oct 2024 · 1 repository · arXiv:2410.10010
-
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs 13 Oct 2024 · 0 repositories · arXiv:2410.12864
-
Joint Mixing Data Augmentation for Skeleton-based Action Recognition 13 Oct 2024 · 1 repository
-
Learning Pattern-Specific Experts for Time Series Forecasting Under Patch-level Distribution Shift 13 Oct 2024 · 1 repository · arXiv:2410.09836
-
Learning to Rank for Multiple Retrieval-Augmented Models through Iterative Utility Maximization 13 Oct 2024 · 0 repositories · arXiv:2410.09942
-
LibEER: A Comprehensive Benchmark and Algorithm Library for EEG-based Emotion Recognition 13 Oct 2024 · 2 repositories · arXiv:2410.09767