Methods › General › Attention Modules
Attention Modules
The archive attaches this collection's text per method and the copies differ: 5 distinct texts across 32 of the 33 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.
Text 1, carried by 26 of 33 methods:
Attention Modules refer to modules that incorporate attention mechanisms. For example, multi-head attention is a module that incorporates multiple attention heads. Below you can find a continuously updating list of attention modules.
Text 2, carried by 3 of 33 methods:
Image Model Blocks are building blocks used in image models such as convolutional neural networks. Below you can find a continuously updating list of image model blocks.
Text 3, carried by 1 of 33 methods:
Convolutions are a type of operation that can be used to learn representations from images. They involve a learnable kernel sliding over the image and performing element-wise multiplication with the input. The specification allows for parameter sharing and translation invariance. Below you can find a continuously updating list of convolutions.
Text 4, carried by 1 of 33 methods:
The original self-attention component in the Transformer architecture has a O(n²) time and memory complexity where n is the input sequence length and thus, is not efficient to scale to long inputs. Attention pattern methods look to reduce this complexity by looking at a subset of the space.
Text 5, carried by 1 of 33 methods:
Pooling Operations are used to pool features together, often downsampling the feature map to a smaller size. They can also induce favourable properties such as translation invariance in image classification, as well as bring together information from different parts of a network in tasks like object detection (e.g. pooling different scales).
Methods
All 33 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Multi-Head Attention | – | 24,855 |
| Graph Self-Attention | – | 64 |
| SCA Semantic Cross Attention | – | 47 |
| Deformable Attention Module | – | 42 |
| LAMA Low-Rank Factorization-based Multi-Head Attention | – | 42 |
| Spatial-Reduction Attention | – | 29 |
| Multi-Head Linear Attention | – | 24 |
| Neighborhood Attention | – | 21 |
| Multi-DConv-Head Attention | – | 15 |
| Global Context Block | – | 12 |
| Triplet Attention | – | 11 |
| Bottleneck Transformer Block | – | 10 |
| Spatial Attention-Guided Mask | – | 7 |
| Mixed Attention Block | – | 6 |
| Single-Headed Attention | – | 6 |
| Channel-wise Cross Attention | – | 5 |
| Feedback Memory | – | 4 |
| Re-Attention Module | – | 4 |
| Attention Free Transformer | – | 3 |
| Attention-augmented Convolution | – | 3 |
| GPSA Gated Positional Self-Attention | – | 3 |
| Hopfield Layer | – | 3 |
| All-Attention Layer | – | 2 |
| CAB Contextual Attention Block | – | 2 |
| LeViT Attention Block | – | 2 |
| Spatially Separable Self-Attention | – | 2 |
| Compact Global Descriptor | – | 1 |
| Cross-Scale Non-Local Attention | – | 1 |
| DeLighT Block | – | 1 |
| MHMA Multi-Heads of Mixed Attention | – | 1 |
| Peer-attention | – | 1 |
| SimAdapter | – | 1 |
| Talking-Heads Attention | – | 1 |