Methods › Natural Language Processing › Autoencoding Transformers

Autoencoding Transformers

19 methods 7,853 papers tagged archive 2025-07-28

The archive attaches this collection's text per method and the copies differ: 3 distinct texts across 19 of the 19 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.

Text 1, carried by 14 of 19 methods:

Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.

Text 2, carried by 4 of 19 methods:

Language Models are models for predicting the next word or character in a document. Below you can find a continuously updating list of language models.

Text 3, carried by 1 of 19 methods:

Làm cho tôi 1 file aim head

Methods

All 19 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

BERT – 6,938
T5 – 708
Electric – 167
DeBERTa – 90
Longformer – 87
XLM – 57
CodeT5 – 32
MobileBERT – 12
ConvBERT – 5
CuBERT – 3
DeeBERT – 3
I-BERT – 3
AutoTinyBERT – 2
DynaBERT – 2
MacBERT – 2
TernaryBERT – 2
BinaryBERT – 1
SMITH Siamese Multi-depth Transformer-based Hierarchical Encoder – 1
SqueezeBERT – 1