Methods › General › Attention Modules

Attention Modules

33 methods 25,082 papers tagged archive 2025-07-28

The archive attaches this collection's text per method and the copies differ: 5 distinct texts across 32 of the 33 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.

Text 1, carried by 26 of 33 methods:

Attention Modules refer to modules that incorporate attention mechanisms. For example, multi-head attention is a module that incorporates multiple attention heads. Below you can find a continuously updating list of attention modules.

Text 2, carried by 3 of 33 methods:

Image Model Blocks are building blocks used in image models such as convolutional neural networks. Below you can find a continuously updating list of image model blocks.

Text 3, carried by 1 of 33 methods:

Convolutions are a type of operation that can be used to learn representations from images. They involve a learnable kernel sliding over the image and performing element-wise multiplication with the input. The specification allows for parameter sharing and translation invariance. Below you can find a continuously updating list of convolutions.

Text 4, carried by 1 of 33 methods:

The original self-attention component in the Transformer architecture has a O(n²) time and memory complexity where n is the input sequence length and thus, is not efficient to scale to long inputs. Attention pattern methods look to reduce this complexity by looking at a subset of the space.

Text 5, carried by 1 of 33 methods:

Pooling Operations are used to pool features together, often downsampling the feature map to a smaller size. They can also induce favourable properties such as translation invariance in image classification, as well as bring together information from different parts of a network in tasks like object detection (e.g. pooling different scales).

Methods

All 33 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

Multi-Head Attention – 24,855
Graph Self-Attention – 64
SCA Semantic Cross Attention – 47
Deformable Attention Module – 42
LAMA Low-Rank Factorization-based Multi-Head Attention – 42
Spatial-Reduction Attention – 29
Multi-Head Linear Attention – 24
Neighborhood Attention – 21
Multi-DConv-Head Attention – 15
Global Context Block – 12
Triplet Attention – 11
Bottleneck Transformer Block – 10
Spatial Attention-Guided Mask – 7
Mixed Attention Block – 6
Single-Headed Attention – 6
Channel-wise Cross Attention – 5
Feedback Memory – 4
Re-Attention Module – 4
Attention Free Transformer – 3
Attention-augmented Convolution – 3
GPSA Gated Positional Self-Attention – 3
Hopfield Layer – 3
All-Attention Layer – 2
CAB Contextual Attention Block – 2
LeViT Attention Block – 2
Spatially Separable Self-Attention – 2
Compact Global Descriptor – 1
Cross-Scale Non-Local Attention – 1
DeLighT Block – 1
MHMA Multi-Heads of Mixed Attention – 1
Peer-attention – 1
SimAdapter – 1
Talking-Heads Attention – 1