Methods › Natural Language Processing › Transformers › DistilBERT
DistilBERT
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
DistilBERT is a small, fast, cheap and light Transformer model based on the BERT architecture. Knowledge distillation is performed during the pre-training phase to reduce the size of a BERT model by 40%. To leverage the inductive biases learned by larger models during pre-training, the authors introduce a triple loss combining language modeling, distillation and cosine-distance losses.
Papers archive 2025-07-28
30 shown of 166, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Enhancing Automatic PT Tagging for MEDLINE Citations Using Transformer-Based Models 3 Jun 2025 · 0 repositories · arXiv:2506.03321
-
Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs 28 May 2025 · 0 repositories · arXiv:2505.22937
-
Less is More: Efficient Weight Farcasting with 1-Layer Neural Network 5 May 2025 · 0 repositories · arXiv:2505.02714
-
Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment 19 Apr 2025 · 0 repositories · arXiv:2504.14232
-
Hybrid Emotion Recognition: Enhancing Customer Interactions Through Acoustic and Textual Analysis 27 Mar 2025 · 0 repositories · arXiv:2503.21927
-
Meme Similarity and Emotion Detection using Multimodal Analysis 21 Mar 2025 · 0 repositories · arXiv:2503.17493
-
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer 20 Mar 2025 · 1 repository · arXiv:2503.16731
-
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer 20 Mar 2025 · 0 repositories · arXiv:2503.15983
-
An End-to-End Homomorphically Encrypted Neural Network 22 Feb 2025 · 0 repositories · arXiv:2502.16176
-
Integrating Language Models for Enhanced Network State Monitoring in DRL-Based SFC Provisioning 16 Feb 2025 · 0 repositories · arXiv:2502.11298
-
Leveraging Conditional Mutual Information to Improve Large Language Model Fine-Tuning For Classification 16 Feb 2025 · 0 repositories · arXiv:2502.11258
-
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions 30 Jan 2025 · 0 repositories · arXiv:2502.12017
-
Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models 14 Jan 2025 · 0 repositories · arXiv:2501.08271
-
Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning 13 Jan 2025 · 0 repositories · arXiv:2501.07663
-
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech 10 Jan 2025 · 0 repositories · arXiv:2501.05755
-
Exploring Variability in Fine-Tuned Models for Text Classification with DistilBERT 31 Dec 2024 · 0 repositories · arXiv:2501.00241
-
Text Classification: Neural Networks VS Machine Learning Models VS Pre-trained Models 30 Dec 2024 · 0 repositories · arXiv:2412.21022
-
Assessing Text Classification Methods for Cyberbullying Detection on Social Media Platforms 27 Dec 2024 · 0 repositories · arXiv:2412.19928
-
Resource-Efficient Transformer Architecture: Optimizing Memory and Execution Time for Real-Time Applications 25 Dec 2024 · 0 repositories · arXiv:2501.00042
-
A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit 23 Nov 2024 · 1 repository · arXiv:2411.15404
-
BERT-Based Approach for Automating Course Articulation Matrix Construction with Explainable AI 21 Nov 2024 · 1 repository · arXiv:2411.14254
-
Autonomous Droplet Microfluidic Design Framework with Large Language Models 11 Nov 2024 · 1 repository · arXiv:2411.06691
-
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm 30 Oct 2024 · 1 repository · arXiv:2410.23182
-
LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model 21 Oct 2024 · 0 repositories · arXiv:2410.15656
-
A Two-Model Approach for Humour Style Recognition 9 Oct 2024 · 1 repository · arXiv:2410.12842
-
A Comparative Study of Hybrid Models in Health Misinformation Text Classification 8 Oct 2024 · 0 repositories · arXiv:2410.06311
-
Depression detection in social media posts using transformer-based models and auxiliary features 30 Sep 2024 · 0 repositories · arXiv:2409.20048
-
Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection 20 Sep 2024 · 0 repositories · arXiv:2409.13331
-
Active Learning to Guide Labeling Efforts for Question Difficulty Estimation 14 Sep 2024 · 1 repository · arXiv:2409.09258
-
Protein sequence classification using natural language processing techniques 6 Sep 2024 · 0 repositories · arXiv:2409.04491
Tasks archive 2025-07-28
20 shown of 172 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections