Attention
The archive attaches this collection's text per method and the copies differ: 2 distinct texts across 11 of the 11 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.
Text 1, carried by 10 of 11 methods:
jinga lala
Text 2, carried by 1 of 11 methods:
Attention Modules refer to modules that incorporate attention mechanisms. For example, multi-head attention is a module that incorporates multiple attention heads. Below you can find a continuously updating list of attention modules.
Methods
All 11 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Class Attention | – | 36 |
| Grouped-query attention | – | 19 |
| Attention Sinks | – | 13 |
| Multi-Query Attention | – | 13 |
| FGA Factor Graph Attention | – | 9 |
| Multi-Attention Network | – | 9 |
| RMN Residual Masking Network | – | 2 |
| Weight excitation | – | 2 |
| CW-CNN & CW-AT CW-Complex Attention Mechanism | – | 1 |
| MHMA Multi-Heads of Mixed Attention | – | 1 |
| Quick Attention | – | 1 |