Methods › Natural Language Processing › Autoencoding Transformers › DeBERTa
DeBERTa
Introduced by Pengcheng He et al. in DeBERTa: Decoding-enhanced BERT with Disentangled Attention
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
DeBERTa is a Transformer-based neural language model that aims to improve the BERT and RoBERTa models with two techniques: a disentangled attention mechanism and an enhanced mask decoder. The disentangled attention mechanism is where each word is represented unchanged using two vectors that encode its content and position, respectively, and the attention weights among words are computed using disentangle matrices on their contents and relative positions. The enhanced mask decoder is used to replace the output softmax layer to predict the masked tokens for model pre-training. In addition, a new virtual adversarial training method is used for fine-tuning to improve model’s generalization on downstream tasks.
Papers archive 2025-07-28
30 shown of 90, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles 15 Jul 2025 · 1 repository · arXiv:2507.11764
-
PlantBert: An Open Source Language Model for Plant Science 10 Jun 2025 · 0 repositories · arXiv:2506.08897
-
WeightLoRA: Keep Only Necessary Adapters 3 Jun 2025 · 0 repositories · arXiv:2506.02724
-
Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking 29 May 2025 · 0 repositories · arXiv:2505.23117
-
Transformer-Based Named Entity Recognition for Automated Server Provisioning 1 Apr 2025 · 1 repository
-
Sarang at DEFACTIFY 4.0: Detecting AI-Generated Text Using Noised Data and an Ensemble of DeBERTa Models 24 Feb 2025 · 0 repositories · arXiv:2502.16857
-
Code-Mixed Telugu-English Hate Speech Detection 15 Feb 2025 · 0 repositories · arXiv:2502.10632
-
Zero-Shot Belief: A Hard Problem for LLMs 12 Feb 2025 · 0 repositories · arXiv:2502.08777
-
Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models 27 Dec 2024 · 0 repositories · arXiv:2412.19449
-
Lightweight Safety Classification Using Pruned Language Models 18 Dec 2024 · 0 repositories · arXiv:2412.13435
-
SEKE: Specialised Experts for Keyword Extraction 18 Dec 2024 · 1 repository · arXiv:2412.14087
-
RAGulator: Lightweight Out-of-Context Detectors for Grounded Text Generation 6 Nov 2024 · 0 repositories · arXiv:2411.03920
-
Bonafide at LegalLens 2024 Shared Task: Using Lightweight DeBERTa Based Encoder For Legal Violation Detection and Resolution 30 Oct 2024 · 1 repository · arXiv:2410.22977
-
Evaluating Transformer Models for Suicide Risk Detection on Social Media 10 Oct 2024 · 0 repositories · arXiv:2410.08375
-
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective 6 Oct 2024 · 1 repository · arXiv:2410.04466
-
Multimodal Coherent Explanation Generation of Robot Failures 1 Oct 2024 · 1 repository · arXiv:2410.00659
-
Improving Academic Skills Assessment with NLP and Ensemble Learning 23 Sep 2024 · 0 repositories · arXiv:2409.19013
-
Instruct-DeBERTa: A Hybrid Approach for Aspect-based Sentiment Analysis on Textual Reviews 23 Aug 2024 · 0 repositories · arXiv:2408.13202
-
Scientific QA System with Verifiable Answers 16 Jul 2024 · 1 repository · arXiv:2407.11485
-
Turn-Level Empathy Prediction Using Psychological Indicators 11 Jul 2024 · 0 repositories · arXiv:2407.08607
-
SecureNet: A Comparative Study of DeBERTa and Large Language Models for Phishing Detection 10 Jun 2024 · 0 repositories · arXiv:2406.06663
-
BERTs are Generative In-Context Learners 7 Jun 2024 · 1 repository · arXiv:2406.04823Syntology ran 13 of 26 samples · 13 unverified
-
Modeling Emotional Trajectories in Written Stories Utilizing Transformers and Weakly-Supervised Learning 4 Jun 2024 · 1 repository · arXiv:2406.02251
-
Explainable Automatic Grading with Neural Additive Models 1 May 2024 · 0 repositories · arXiv:2405.00489
-
Comparative Analysis of Deep Natural Networks and Large Language Models for Aspect-Based Sentiment Analysis 17 Apr 2024 · 1 repository
-
DKE-Research at SemEval-2024 Task 2: Incorporating Data Augmentation with Generative Models and Biomedical Knowledge to Enhance Inference Robustness 14 Apr 2024 · 0 repositories · arXiv:2404.09206
-
TLDR at SemEval-2024 Task 2: T5-generated clinical-Language summaries for DeBERTa Report Analysis 14 Apr 2024 · 1 repository · arXiv:2404.09136
-
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining 29 Jan 2024 · 0 repositories · arXiv:2401.15861
-
Enhancing Essay Scoring with Adversarial Weights Perturbation and Metric-specific AttentionPooling 6 Jan 2024 · 0 repositories · arXiv:2401.05433
-
More than Correlation: Do Large Language Models Learn Causal Representations of Space? 26 Dec 2023 · 0 repositories · arXiv:2312.16257
Tasks archive 2025-07-28
20 shown of 133 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections