Methods › Natural Language Processing › Attention Patterns › Routing Attention
Routing Attention
Introduced by Aurko Roy et al. in Efficient Content-Based Sparse Attention with Routing Transformers
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Routed Attention is an attention pattern proposed as part of the Routing Transformer architecture. Each attention module considers a clustering of the space: the current timestep only attends to context belonging to the same cluster. In other word, the current time-step query is routed to a limited number of context through its cluster assignment. This can be contrasted with strided attention patterns and those proposed with the Sparse Transformer.
In the image to the right, the rows represent the outputs while the columns represent the inputs. The different colors represent cluster memberships for the output token.
Papers archive 2025-07-28
13 shown of 13, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Lightweight Relational Embedding in Task-Interpolated Few-Shot Networks for Enhanced Gastrointestinal Disease Classification 30 May 2025 · 0 repositories · arXiv:2505.24792
-
CL-MFAP: A Contrastive Learning-Based Multimodal Foundation Model for Molecular Property Prediction and Antibiotic Screening 16 Feb 2025 · 1 repository · arXiv:2502.11001Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)
-
HYATT-Net is Grand: A Hybrid Attention Network for Performant Anatomical Landmark Detection 9 Dec 2024 · 1 repository · arXiv:2412.06499
-
Pubic Symphysis-Fetal Head Segmentation Network Using BiFormer Attention Mechanism and Multipath Dilated Convolution 14 Oct 2024 · 0 repositories · arXiv:2410.10352
-
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention 11 Oct 2024 · 1 repository · arXiv:2410.08582
-
Vision Transformer with Key-select Routing Attention for Single Image Dehazing 28 Jun 2024 · 0 repositories · arXiv:2406.19703
-
BRAU-Net++: U-Shaped Hybrid CNN-Transformer Network for Medical Image Segmentation 1 Jan 2024 · 1 repository · arXiv:2401.00722
-
Pubic Symphysis-Fetal Head Segmentation Using Pure Transformer with Bi-level Routing Attention 30 Sep 2023 · 0 repositories · arXiv:2310.00289
-
BGF-YOLO: Enhanced YOLOv8 with Multiscale Attentional Feature Fusion for Brain Tumor Detection 22 Sep 2023 · 1 repository · arXiv:2309.12585
-
Student Classroom Behavior Detection based on YOLOv7-BRA and Multi-Model Fusion 13 May 2023 · 1 repository · arXiv:2305.07825
-
Hybrid Routing Transformer for Zero-Shot Learning 29 Mar 2022 · 0 repositories · arXiv:2203.15310
-
Hurdles to Progress in Long-form Question Answering 10 Mar 2021 · 2 repositories · arXiv:2103.06332
-
Efficient Content-Based Sparse Attention with Routing Transformers 12 Mar 2020 · 2 repositories · arXiv:2003.05997Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 37 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections