Methods › General › Attention Mechanisms

Attention Mechanisms

62 methods 32,081 papers tagged archive 2025-07-28

The archive attaches this collection's text per method and the copies differ: 3 distinct texts across 3 of the 62 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.

Text 1, carried by 1 of 62 methods:

Attention Modules refer to modules that incorporate attention mechanisms. For example, multi-head attention is a module that incorporates multiple attention heads. Below you can find a continuously updating list of attention modules.

Text 2, carried by 1 of 62 methods:

The original self-attention component in the Transformer architecture has a O(n²) time and memory complexity where n is the input sequence length and thus, is not efficient to scale to long inputs. Attention pattern methods look to reduce this complexity by looking at a subset of the space.

Text 3, carried by 1 of 62 methods:

jinga lala

Also reached at /methods/category/attention-mechanisms-1 (Papers with Code's slug for this collection; the archive carries no slugs, so this site's is derived from the name).

Methods

All 62 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

Attention Attention Is All You Need – 31,583
FAVOR+ Fast Attention Via Positive Orthogonal Random Features – 102
Global-Local Attention – 56
DMA Dual Multimodal Attention – 39
Location-based Attention – 37
Class Attention – 36
BAM Bottleneck Attention Module – 33
Content-based Attention – 33
SRM style-based recalibration module – 33
TAM Temporal Adaptive Module – 32
Adaptive Masking – 31
Coordinate attention – 31
ECANet efficient channel attention – 29
Location Sensitive Attention – 25
Highway networks – 24
LSH Attention Locality Sensitive Hashing Attention – 22
Neighborhood Attention – 21
Set Transformer – 20
SEAM Self-supervised Equivariant Attention Mechanism – 14
Bi-attention Bilinear Attention – 11
RGA Relation-aware Global Attention – 11
DANet Dual Attention Network – 10
GALA Global-and-Local attention – 9
GCT Gated Channel Transformation – 8
Cross-Covariance Attention – 7
Multiplicative Attention – 6
SPNet Strip Pooling Network – 5
scSE Spatial and Channel SE Blocks – 4
3D SA 3 Dimensional Soft Attention – 3
Branch attention – 3
Channel Squeeze and Spatial Excitation Channel Squeeze and Spatial Excitation (sSE) – 3
Spatial & Temporal Attention – 3
Channel & Spatial attention – 2
Class Activation Guided Attention Mechanism Class Activation Guided Attention Mechanism (CAGAM) – 2
Concurrent Spatial and Channel Squeeze & Excitation Concurrent Spatial and Channel Squeeze & Excitation (scSE) – 2
Deformable ConvNets Deformable Convolutional Networks – 2
FcaNet Frequency channel attention networks – 2
Global Sub-Sampled Attention – 2
Locally-Grouped Self-Attention – 2
SCA-CNN Spatial and Channel-wise Attention-based Convolutional Neural Network – 2
STA-LSTM Spatio-Temporal Attention LSTM – 2
Self-Calibrated Convolutions – 2
Sparse Sinkhorn Attention – 2
Attention Feature Filters – 1
Dense Synthesized Attention – 1
Factorized Dense Synthesized Attention – 1
Factorized Random Synthesized Attention – 1
Fast Voxel Query – 1
GSoP-Net Global second-order pooling convolutional networks – 1
Gather-Excite Networks – 1
HyperHyperNetwork Hyper HyperNetwork – 1
HyperSA HyperGraph Self-Attention – 1
MHMA Multi-Heads of Mixed Attention – 1
Multidimensional Convolutional Frequency Attention – 1
ProCAN Progressive Growing Channel Attentive Non-Local Network – 1
Random Synthesized Attention – 1
Segregated Attention Network – 1
SortCut Sinkhorn Attention – 1
Vision Eagle Attention – 1
WERSA Wavelet-Enhanced Random Spectral Attention – 1
Differential attention for visual question answering – 0
Spatial Attention ConvMixer – 0