Methods › Natural Language Processing › Autoencoding Transformers › DeBERTa

DeBERTa

90 papers tagged archive 2025-07-28

Introduced by Pengcheng He et al. in DeBERTa: Decoding-enhanced BERT with Disentangled Attention

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

DeBERTa is a Transformer-based neural language model that aims to improve the BERT and RoBERTa models with two techniques: a disentangled attention mechanism and an enhanced mask decoder. The disentangled attention mechanism is where each word is represented unchanged using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangle matrices on their contents and relative positions. The enhanced mask decoder is used to replace the output softmax layer to predict the masked tokens for model pre-training. In addition, a new virtual adversarial training method is used for fine-tuning to improve model’s generalization on downstream tasks.

PaperSource

Papers archive 2025-07-28

30 shown of 90, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 133 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling25
Language Modeling19
Sentence11
Natural Language Inference9
Question Answering9
Natural Language Understanding8
Sentiment Analysis8
Large Language Model6
Aspect-Based Sentiment Analysis5
Aspect-Based Sentiment Analysis (ABSA)5
Named Entity Recognition (NER)5
Articles4
Data Augmentation4
Hate Speech Detection4
Self-Supervised Learning4
Text Classification4
Text Generation4
text-classification4
Benchmarking3
Classification3

Usage over time archive 2025-07-28

Papers per year tagged with DeBERTa: 2020 to 2025, peak 29 29 0 2020: 1 paper 2020 2021: 8 papers 2021 2022: 23 papers 2022 2023: 29 papers 2023 2024: 21 papers 2024 2025: 8 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (90 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Autoencoding TransformersTransformers

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections