Methods › Natural Language Processing › Autoregressive Transformers
Autoregressive Transformers
Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.
Methods
All 13 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Transformer | – | 13,999 |
| GPT | – | 1,212 |
| GPT-2 | – | 768 |
| Transformer-XL | – | 64 |
| Universal Transformer | – | 18 |
| Linformer | – | 17 |
| Primer | – | 14 |
| Levenshtein Transformer | – | 12 |
| Routing Transformer | – | 3 |
| Feedback Transformer | – | 2 |
| Sandwich Transformer | – | 2 |
| DeLighT | – | 1 |
| Sinkhorn Transformer | – | 1 |