{"url":"/method/deberta","slug":"deberta","name":"DeBERTa","full_name":"DeBERTa","full_name_withheld":false,"description_markdown":"**DeBERTa** is a [Transformer](https://paperswithcode.com/methods/category/transformers)-based neural language model that aims to improve the [BERT](https://paperswithcode.com/method/bert) and [RoBERTa](https://paperswithcode.com/method/roberta) models with two techniques: a [disentangled attention mechanism](https://paperswithcode.com/method/disentangled-attention-mechanism) and an enhanced mask decoder. The disentangled attention mechanism is where each word is represented unchanged using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangle matrices on their contents and relative positions. The enhanced mask decoder is used to replace the output [softmax](https://paperswithcode.com/method/softmax) layer to predict the masked tokens for model pre-training.  In addition, a new virtual adversarial training method is used for fine-tuning to improve model’s generalization on downstream tasks.","description_state":"present","introduced_year":null,"introduced_by":{"title":"DeBERTa: Decoding-enhanced BERT with Disentangled Attention","paper":"/paper/deberta-decoding-enhanced-bert-with","first_author":"Pengcheng He","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/deberta-decoding-enhanced-bert-with"},"source":{"url":"https://arxiv.org/abs/2006.03654v6","title":"DeBERTa: Decoding-enhanced BERT with Disentangled Attention","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":90,"archive_num_papers":90,"papers_newest_first":[{"paper":"/paper/ai-wizards-at-checkthat-2025-enhancing","title":"AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles","date":"2025-07-15","arxiv_id":"2507.11764","n_code_links":1,"syntology":null},{"paper":null,"title":"PlantBert: An Open Source Language Model for Plant Science","date":"2025-06-10","arxiv_id":"2506.08897","n_code_links":0,"syntology":null},{"paper":null,"title":"WeightLoRA: Keep Only Necessary Adapters","date":"2025-06-03","arxiv_id":"2506.02724","n_code_links":0,"syntology":null},{"paper":null,"title":"Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking","date":"2025-05-29","arxiv_id":"2505.23117","n_code_links":0,"syntology":null},{"paper":"/paper/transformer-based-named-entity-recognition-2","title":"Transformer-Based Named Entity Recognition for Automated Server Provisioning","date":"2025-04-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models","date":"2025-02-24","arxiv_id":"2502.16857","n_code_links":0,"syntology":null},{"paper":null,"title":"Code-Mixed Telugu-English Hate Speech Detection","date":"2025-02-15","arxiv_id":"2502.10632","n_code_links":0,"syntology":null},{"paper":null,"title":"Zero-Shot Belief: A Hard Problem for LLMs","date":"2025-02-12","arxiv_id":"2502.08777","n_code_links":0,"syntology":null},{"paper":null,"title":"Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models","date":"2024-12-27","arxiv_id":"2412.19449","n_code_links":0,"syntology":null},{"paper":null,"title":"Lightweight Safety Classification Using Pruned Language Models","date":"2024-12-18","arxiv_id":"2412.13435","n_code_links":0,"syntology":null},{"paper":"/paper/seke-specialised-experts-for-keyword","title":"SEKE: Specialised Experts for Keyword Extraction","date":"2024-12-18","arxiv_id":"2412.14087","n_code_links":1,"syntology":null},{"paper":null,"title":"RAGulator: Lightweight Out-of-Context Detectors for Grounded Text Generation","date":"2024-11-06","arxiv_id":"2411.03920","n_code_links":0,"syntology":null},{"paper":"/paper/bonafide-at-legallens-2024-shared-task-using","title":"Bonafide at LegalLens 2024 Shared Task: Using Lightweight DeBERTa Based Encoder For Legal Violation Detection and Resolution","date":"2024-10-30","arxiv_id":"2410.22977","n_code_links":1,"syntology":null},{"paper":null,"title":"Evaluating Transformer Models for Suicide Risk Detection on Social Media","date":"2024-10-10","arxiv_id":"2410.08375","n_code_links":0,"syntology":null},{"paper":"/paper/large-language-model-inference-acceleration-a","title":"Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective","date":"2024-10-06","arxiv_id":"2410.04466","n_code_links":1,"syntology":null},{"paper":"/paper/multimodal-coherent-explanation-generation-of","title":"Multimodal Coherent Explanation Generation of Robot Failures","date":"2024-10-01","arxiv_id":"2410.00659","n_code_links":1,"syntology":null},{"paper":null,"title":"Improving Academic Skills Assessment with NLP and Ensemble Learning","date":"2024-09-23","arxiv_id":"2409.19013","n_code_links":0,"syntology":null},{"paper":null,"title":"Instruct-DeBERTa: A Hybrid Approach for Aspect-based Sentiment Analysis on Textual Reviews","date":"2024-08-23","arxiv_id":"2408.13202","n_code_links":0,"syntology":null},{"paper":"/paper/scientific-qa-system-with-verifiable-answers","title":"Scientific QA System with Verifiable Answers","date":"2024-07-16","arxiv_id":"2407.11485","n_code_links":1,"syntology":null},{"paper":null,"title":"Turn-Level Empathy Prediction Using Psychological Indicators","date":"2024-07-11","arxiv_id":"2407.08607","n_code_links":0,"syntology":null},{"paper":null,"title":"SecureNet: A Comparative Study of DeBERTa and Large Language Models for Phishing Detection","date":"2024-06-10","arxiv_id":"2406.06663","n_code_links":0,"syntology":null},{"paper":"/paper/berts-are-generative-in-context-learners","title":"BERTs are Generative In-Context Learners","date":"2024-06-07","arxiv_id":"2406.04823","n_code_links":1,"syntology":{"ran":13,"of":26,"unverified":13,"pointer_only":0}},{"paper":"/paper/modeling-emotional-trajectories-in-written","title":"Modeling Emotional Trajectories in Written Stories Utilizing Transformers and Weakly-Supervised Learning","date":"2024-06-04","arxiv_id":"2406.02251","n_code_links":1,"syntology":null},{"paper":null,"title":"Explainable Automatic Grading with Neural Additive Models","date":"2024-05-01","arxiv_id":"2405.00489","n_code_links":0,"syntology":null},{"paper":"/paper/comparative-analysis-of-deep-natural-networks","title":"Comparative Analysis of Deep Natural Networks and Large Language Models for Aspect-Based Sentiment Analysis","date":"2024-04-17","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"DKE-Research at SemEval-2024 Task 2: Incorporating Data Augmentation with Generative Models and Biomedical Knowledge to Enhance Inference Robustness","date":"2024-04-14","arxiv_id":"2404.09206","n_code_links":0,"syntology":null},{"paper":"/paper/tldr-at-semeval-2024-task-2-t5-generated","title":"TLDR at SemEval-2024 Task 2: T5-generated clinical-Language summaries for DeBERTa Report Analysis","date":"2024-04-14","arxiv_id":"2404.09136","n_code_links":1,"syntology":null},{"paper":null,"title":"BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining","date":"2024-01-29","arxiv_id":"2401.15861","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling","date":"2024-01-06","arxiv_id":"2401.05433","n_code_links":0,"syntology":null},{"paper":null,"title":"More than Correlation: Do Large Language Models Learn Causal Representations of Space?","date":"2023-12-26","arxiv_id":"2312.16257","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":25},{"task":"/task/language-modeling","name":"Language Modeling","papers":19},{"task":"/task/sentence","name":"Sentence","papers":11},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":9},{"task":"/task/question-answering","name":"Question Answering","papers":9},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":8},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":8},{"task":"/task/large-language-model","name":"Large Language Model","papers":6},{"task":"/task/aspect-based-sentiment-analysis-1","name":"Aspect-Based Sentiment Analysis","papers":5},{"task":"/task/aspect-based-sentiment-analysis","name":"Aspect-Based Sentiment Analysis (ABSA)","papers":5},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":5},{"task":"/task/articles","name":"Articles","papers":4},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":4},{"task":"/task/hate-speech-detection","name":"Hate Speech Detection","papers":4},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":4},{"task":"/task/text-classification","name":"Text Classification","papers":4},{"task":"/task/text-generation","name":"Text Generation","papers":4},{"task":"/task/text-classification-1","name":"text-classification","papers":4},{"task":"/task/benchmarking","name":"Benchmarking","papers":3},{"task":"/task/classification-1","name":"Classification","papers":3}],"tasks_shown":20,"n_tasks":133,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":8},{"year":"2022","papers":23},{"year":"2023","papers":29},{"year":"2024","papers":21},{"year":"2025","papers":8}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/deberta"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}