Methods › Natural Language Processing › Language Models
Language Models
The archive attaches this collection's text per method and the copies differ: 5 distinct texts across 49 of the 51 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.
Text 1, carried by 42 of 51 methods:
Language Models are models for predicting the next word or character in a document. Below you can find a continuously updating list of language models.
Text 2, carried by 4 of 51 methods:
Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.
Text 3, carried by 1 of 51 methods:
Interpretability Methods seek to explain the predictions made by neural networks by introducing mechanisms to enduce or enforce interpretability. For example, LIME approximates the neural network with a locally interpretable model. Below you can find a continuously updating list of interpretability methods.
Text 4, carried by 1 of 51 methods:
Generative Models aim to model data generatively (rather than discriminatively), that is they aim to approximate the probability distribution of the data. Below you can find a continuously updating list of generative models for computer vision.
Text 5, carried by 1 of 51 methods:
Fine-Tuning methods in deep learning take existing trained networks and 'fine-tune' them to a new task so that information contained in the weights can be repurposed. Below you can find a continuously updating list of fine-tuning methods.
Methods
All 51 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Diffusion | – | 13,848 |
| BERT | – | 6,938 |
| GPT-4 | – | 2,870 |
| LLaMA | – | 1,062 |
| OPT | – | 285 |
| CAM Class-activation map | – | 267 |
| ELMo | – | 234 |
| mBERT | – | 198 |
| XLM-R | – | 176 |
| Electric | – | 167 |
| BLOOM | – | 116 |
| Flan-T5 | – | 107 |
| mT5 | – | 104 |
| mBART | – | 66 |
| Pythia | – | 60 |
| GLM | – | 46 |
| ULMFiT Universal Language Model Fine-tuning | – | 40 |
| Synthesizer | – | 29 |
| BLOOMZ | – | 26 |
| Chinchilla | – | 25 |
| CodeGen | – | 25 |
| GEE Generative Emotion Estimator | – | 12 |
| GPT-NeoX | – | 11 |
| ProphetNet | – | 11 |
| Galactica | – | 10 |
| CharacterBERT | – | 8 |
| UL2 | – | 8 |
| mT0 | – | 8 |
| CANINE | – | 6 |
| Bort | – | 5 |
| Neural Probabilistic Language Model | 2003 | 5 |
| Cross-encoder Reranking | – | 4 |
| CuBERT | – | 3 |
| Gated Convolution Network | – | 3 |
| OPT-IML | – | 3 |
| PMLM Probabilistically Masked Language Model | – | 3 |
| DynaBERT | – | 2 |
| Feedback Transformer | – | 2 |
| SHA-RNN Single Headed Attention RNN | – | 2 |
| Sandwich Transformer | – | 2 |
| Step-DPO Step-wise Direct Preference Optimization | – | 2 |
| CPM-2 | – | 1 |
| DAHSF Digestion Algorithm in Hierarchical Symbolic Forests | – | 1 |
| DeLighT | – | 1 |
| ERNIE-GEN | – | 1 |
| KE-MLM Knowledge Enhanced Masked Language Model | – | 1 |
| MixLoRA | – | 1 |
| PanGu-α | – | 1 |
| mBARTHez | – | 1 |
| ncMUFCMU | – | 1 |
| ooJpiued | – | 1 |