Methods › General › Attention Mechanisms › Sparse Sinkhorn Attention
Sparse Sinkhorn Attention
Introduced by Yi Tay et al. in Sparse Sinkhorn Attention
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Sparse Sinkhorn Attention is an attention mechanism that reduces the memory complexity of the dot-product attention mechanism and is capable of learning sparse attention outputs. It is based on the idea of differentiable sorting of internal representations within the self-attention module. SSA incorporates a meta sorting network that learns to rearrange and sort input sequences. Sinkhorn normalization is used to normalize the rows and columns of the sorting matrix. The actual SSA attention mechanism then acts on the block sorted sequences.
Papers archive 2025-07-28
2 shown of 2, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
VTP: Volumetric Transformer for Multi-view Multi-person 3D Pose Estimation 25 May 2022 · 0 repositories · arXiv:2205.12602
-
Sparse Sinkhorn Attention 26 Feb 2020 · 1 repository · arXiv:2002.11296
Tasks archive 2025-07-28
9 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections