Methods › Natural Language Processing › Language Models

Language Models

51 methods 25,913 papers tagged archive 2025-07-28

The archive attaches this collection's text per method and the copies differ: 5 distinct texts across 49 of the 51 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.

Text 1, carried by 42 of 51 methods:

Language Models are models for predicting the next word or character in a document. Below you can find a continuously updating list of language models.

Text 2, carried by 4 of 51 methods:

Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.

Text 3, carried by 1 of 51 methods:

Interpretability Methods seek to explain the predictions made by neural networks by introducing mechanisms to enduce or enforce interpretability. For example, LIME approximates the neural network with a locally interpretable model. Below you can find a continuously updating list of interpretability methods.

Text 4, carried by 1 of 51 methods:

Generative Models aim to model data generatively (rather than discriminatively), that is they aim to approximate the probability distribution of the data. Below you can find a continuously updating list of generative models for computer vision.

Text 5, carried by 1 of 51 methods:

Fine-Tuning methods in deep learning take existing trained networks and 'fine-tune' them to a new task so that information contained in the weights can be repurposed. Below you can find a continuously updating list of fine-tuning methods.

Methods

All 51 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

Diffusion – 13,848
BERT – 6,938
GPT-4 – 2,870
LLaMA – 1,062
OPT – 285
CAM Class-activation map – 267
ELMo – 234
mBERT – 198
XLM-R – 176
Electric – 167
BLOOM – 116
Flan-T5 – 107
mT5 – 104
mBART – 66
Pythia – 60
GLM – 46
ULMFiT Universal Language Model Fine-tuning – 40
Synthesizer – 29
BLOOMZ – 26
Chinchilla – 25
CodeGen – 25
GEE Generative Emotion Estimator – 12
GPT-NeoX – 11
ProphetNet – 11
Galactica – 10
CharacterBERT – 8
UL2 – 8
mT0 – 8
CANINE – 6
Bort – 5
Neural Probabilistic Language Model 2003 5
Cross-encoder Reranking – 4
CuBERT – 3
Gated Convolution Network – 3
OPT-IML – 3
PMLM Probabilistically Masked Language Model – 3
DynaBERT – 2
Feedback Transformer – 2
SHA-RNN Single Headed Attention RNN – 2
Sandwich Transformer – 2
Step-DPO Step-wise Direct Preference Optimization – 2
CPM-2 – 1
DAHSF Digestion Algorithm in Hierarchical Symbolic Forests – 1
DeLighT – 1
ERNIE-GEN – 1
KE-MLM Knowledge Enhanced Masked Language Model – 1
MixLoRA – 1
PanGu-α – 1
mBARTHez – 1
ncMUFCMU – 1
ooJpiued – 1