{"url":"/method/distilbert","slug":"distilbert","name":"DistilBERT","full_name":"DistilBERT","full_name_withheld":false,"description_markdown":"**DistilBERT**  is a small, fast, cheap and light [Transformer](https://paperswithcode.com/method/transformer) model based on the [BERT](https://paperswithcode.com/method/bert) architecture. Knowledge distillation is performed during the pre-training phase to reduce the size of a BERT model by 40%. To leverage the inductive biases learned by larger models during pre-training, the authors introduce a triple loss combining language modeling, distillation and cosine-distance losses.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1910.01108v4","title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":166,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Enhancing Automatic PT Tagging for MEDLINE Citations Using Transformer-Based Models","date":"2025-06-03","arxiv_id":"2506.03321","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs","date":"2025-05-28","arxiv_id":"2505.22937","n_code_links":0,"syntology":null},{"paper":null,"title":"Less is More: Efficient Weight Farcasting with 1-Layer Neural Network","date":"2025-05-05","arxiv_id":"2505.02714","n_code_links":0,"syntology":null},{"paper":null,"title":"Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment","date":"2025-04-19","arxiv_id":"2504.14232","n_code_links":0,"syntology":null},{"paper":null,"title":"Hybrid Emotion Recognition: Enhancing Customer Interactions Through Acoustic and Textual Analysis","date":"2025-03-27","arxiv_id":"2503.21927","n_code_links":0,"syntology":null},{"paper":null,"title":"Meme Similarity and Emotion Detection using Multimodal Analysis","date":"2025-03-21","arxiv_id":"2503.17493","n_code_links":0,"syntology":null},{"paper":"/paper/design-and-implementation-of-an-fpga-based","title":"Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer","date":"2025-03-20","arxiv_id":"2503.16731","n_code_links":1,"syntology":null},{"paper":null,"title":"InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer","date":"2025-03-20","arxiv_id":"2503.15983","n_code_links":0,"syntology":null},{"paper":null,"title":"An End-to-End Homomorphically Encrypted Neural Network","date":"2025-02-22","arxiv_id":"2502.16176","n_code_links":0,"syntology":null},{"paper":null,"title":"Integrating Language Models for Enhanced Network State Monitoring in DRL-Based SFC Provisioning","date":"2025-02-16","arxiv_id":"2502.11298","n_code_links":0,"syntology":null},{"paper":null,"title":"Leveraging Conditional Mutual Information to Improve Large Language Model Fine-Tuning For Classification","date":"2025-02-16","arxiv_id":"2502.11258","n_code_links":0,"syntology":null},{"paper":null,"title":"Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions","date":"2025-01-30","arxiv_id":"2502.12017","n_code_links":0,"syntology":null},{"paper":null,"title":"Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models","date":"2025-01-14","arxiv_id":"2501.08271","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning","date":"2025-01-13","arxiv_id":"2501.07663","n_code_links":0,"syntology":null},{"paper":null,"title":"CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech","date":"2025-01-10","arxiv_id":"2501.05755","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Variability in Fine-Tuned Models for Text Classification with DistilBERT","date":"2024-12-31","arxiv_id":"2501.00241","n_code_links":0,"syntology":null},{"paper":null,"title":"Text Classification: Neural Networks VS Machine Learning Models VS Pre-trained Models","date":"2024-12-30","arxiv_id":"2412.21022","n_code_links":0,"syntology":null},{"paper":null,"title":"Assessing Text Classification Methods for Cyberbullying Detection on Social Media Platforms","date":"2024-12-27","arxiv_id":"2412.19928","n_code_links":0,"syntology":null},{"paper":null,"title":"Resource-Efficient Transformer Architecture: Optimizing Memory and Execution Time for Real-Time Applications","date":"2024-12-25","arxiv_id":"2501.00042","n_code_links":0,"syntology":null},{"paper":"/paper/a-comparative-analysis-of-transformer-and","title":"A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit","date":"2024-11-23","arxiv_id":"2411.15404","n_code_links":1,"syntology":null},{"paper":"/paper/bert-based-approach-for-automating-course","title":"BERT-Based Approach for Automating Course Articulation Matrix Construction with Explainable AI","date":"2024-11-21","arxiv_id":"2411.14254","n_code_links":1,"syntology":null},{"paper":"/paper/autonomous-droplet-microfluidic-design","title":"Autonomous Droplet Microfluidic Design Framework with Large Language Models","date":"2024-11-11","arxiv_id":"2411.06691","n_code_links":1,"syntology":null},{"paper":"/paper/protransformer-robustify-transformers-via","title":"ProTransformer: Robustify Transformers via Plug-and-Play Paradigm","date":"2024-10-30","arxiv_id":"2410.23182","n_code_links":1,"syntology":null},{"paper":null,"title":"LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model","date":"2024-10-21","arxiv_id":"2410.15656","n_code_links":0,"syntology":null},{"paper":"/paper/a-two-model-approach-for-humour-style","title":"A Two-Model Approach for Humour Style Recognition","date":"2024-10-09","arxiv_id":"2410.12842","n_code_links":1,"syntology":null},{"paper":null,"title":"A Comparative Study of Hybrid Models in Health Misinformation Text Classification","date":"2024-10-08","arxiv_id":"2410.06311","n_code_links":0,"syntology":null},{"paper":null,"title":"Depression detection in social media posts using transformer-based models and auxiliary features","date":"2024-09-30","arxiv_id":"2409.20048","n_code_links":0,"syntology":null},{"paper":null,"title":"Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection","date":"2024-09-20","arxiv_id":"2409.13331","n_code_links":0,"syntology":null},{"paper":"/paper/active-learning-to-guide-labeling-efforts-for","title":"Active Learning to Guide Labeling Efforts for Question Difficulty Estimation","date":"2024-09-14","arxiv_id":"2409.09258","n_code_links":1,"syntology":null},{"paper":null,"title":"Protein sequence classification using natural language processing techniques","date":"2024-09-06","arxiv_id":"2409.04491","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":23},{"task":"/task/text-classification","name":"Text Classification","papers":22},{"task":"/task/classification-1","name":"Classification","papers":21},{"task":"/task/text-classification-1","name":"text-classification","papers":20},{"task":"/task/language-modelling","name":"Language Modelling","papers":18},{"task":"/task/language-modeling","name":"Language Modeling","papers":14},{"task":"/task/question-answering","name":"Question Answering","papers":14},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":12},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":9},{"task":"/task/sentence","name":"Sentence","papers":9},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":9},{"task":"/task/quantization","name":"Quantization","papers":8},{"task":"/task/model-compression","name":"Model Compression","papers":7},{"task":"/task/classification","name":"General Classification","papers":6},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":6},{"task":"/task/sentiment-classification","name":"Sentiment Classification","papers":6},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":6},{"task":"/task/model","name":"model","papers":6},{"task":null,"name":"GPU","papers":5},{"task":"/task/large-language-model","name":"Large Language Model","papers":5}],"tasks_shown":20,"n_tasks":172,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":27},{"year":"2021","papers":34},{"year":"2022","papers":22},{"year":"2023","papers":34},{"year":"2024","papers":33},{"year":"2025","papers":15}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/distilbert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}