Methods › Natural Language Processing › Transformers

Transformers

85 methods 38,673 papers tagged archive 2025-07-28

The archive attaches this collection's text per method and the copies differ: 5 distinct texts across 83 of the 85 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.

Text 1, carried by 78 of 85 methods:

Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.

Text 2, carried by 2 of 85 methods:

Language Models are models for predicting the next word or character in a document. Below you can find a continuously updating list of language models.

Text 3, carried by 1 of 85 methods:

Fine-Tuning methods in deep learning take existing trained networks and 'fine-tune' them to a new task so that information contained in the weights can be repurposed. Below you can find a continuously updating list of fine-tuning methods.

Text 4, carried by 1 of 85 methods:

Attention Modules refer to modules that incorporate attention mechanisms. For example, multi-head attention is a module that incorporates multiple attention heads. Below you can find a continuously updating list of attention modules.

Text 5, carried by 1 of 85 methods:

Làm cho tôi 1 file aim head

Methods

All 85 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

Focus – 15,340
Transformer – 13,999
BERT – 6,938
GPT-3 – 1,906
BART – 1,642
RAG – 1,286
GPT – 1,212
RoBERTa – 913
GPT-2 – 768
T5 – 708
ALBERT – 172
Electric – 167
XLNet – 167
DistilBERT – 166
PaLM Pathways Language Model – 151
ELECTRA – 118
Performer – 103
DeBERTa – 90
Longformer – 87
CodeBERT – 66
Transformer-XL – 64
XLM – 57
ERNIE – 54
PEGASUS – 53
Sparse Transformer – 47
GPT-Neo – 38
ETC Extended Transformer Construction – 36
CodeT5 – 32
ViLBERT Vision-and-Language BERT – 30
CTRL – 28
Reformer – 20
Switch Transformer – 20
Universal Transformer – 18
Linformer – 17
BigBird – 16
MATE – 14
Primer – 14
E-Branchformer – 13
Levenshtein Transformer – 12
MobileBERT – 12
TNT Transformer in Transformer – 12
ProphetNet – 11
EGT Edge-augmented Graph Transformer – 10
PLATO-2 – 6
Charformer – 5
ConvBERT – 5
Fastformer – 4
Adaptive Span Transformer – 3
Compressive Transformer – 3
DeeBERT – 3
GANformer Generative Adversarial Transformer – 3
I-BERT – 3
Parallel Layers – 3
Routing Transformer – 3
Subformer – 3
TAPEX Table Pre-training via Execution – 3
VideoBERT – 3
AutoTinyBERT – 2
ESACL Enhanced Seq2Seq Autoencoder via Contrastive Learning – 2
Feedback Transformer – 2
Funnel Transformer – 2
MacBERT – 2
Nyströmformer – 2
Sandwich Transformer – 2
TernaryBERT – 2
Adaptively Sparse Transformer – 1
BP-Transformer – 1
BinaryBERT – 1
Chinese Pre-trained Unbalanced Transformer – 1
ClipBERT – 1
DeLighT – 1
IB-BERT Inverted Bottleneck BERT – 1
MHMA Multi-Heads of Mixed Attention – 1
MixLoRA – 1
NormFormer – 1
PAR Transformer – 1
PermuteFormer – 1
PolyNorm Polynomial Composition Activations – 1
RealFormer – 1
SC-GPT – 1
SMITH Siamese Multi-depth Transformer-based Hierarchical Encoder – 1
Sinkhorn Transformer – 1
SongNet – 1
SqueezeBERT – 1
T-D Transformer Decoder – 1