{"url":"/method/knowledge-distillation","slug":"knowledge-distillation","name":"Knowledge Distillation","full_name":"Knowledge Distillation","full_name_withheld":false,"description_markdown":"A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions. Unfortunately, making predictions using a whole ensemble of models is cumbersome and may be too computationally expensive to allow deployment to a large number of users, especially if the individual models are large neural nets. Caruana and his collaborators have shown that it is possible to compress the knowledge in an ensemble into a single model which is much easier to deploy and we develop this approach further using a different compression technique. We achieve some surprising results on MNIST and we show that we can significantly improve the acoustic model of a heavily used commercial system by distilling the knowledge in an ensemble of models into a single model. We also introduce a new type of ensemble composed of one or more full models and many specialist models which learn to distinguish fine-grained classes that the full models confuse. Unlike a mixture of experts, these specialist models can be trained rapidly and in parallel.\r\nSource: [Distilling the Knowledge in a Neural Network](https://arxiv.org/abs/1503.02531)","description_state":"present","introduced_year":null,"introduced_by":{"title":"Distilling the Knowledge in a Neural Network","paper":"/paper/distilling-the-knowledge-in-a-neural-network","first_author":"Geoffrey Hinton","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/distilling-the-knowledge-in-a-neural-network"},"source":{"url":"http://arxiv.org/abs/1503.02531v1","title":"Distilling the Knowledge in a Neural Network","url_on_a_paper_host":true},"code_snippet_url":"https://research.google/blog/auto-generated-summaries-in-google-docs/","code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Knowledge Distillation","url":"/methods/category/knowledge-distillation","pwc_aliases":[]}],"n_papers_tagged":3071,"archive_num_papers":3071,"papers_newest_first":[{"paper":"/paper/dvfl-net-a-lightweight-distilled-video-focal","title":"DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition","date":"2025-07-16","arxiv_id":"2507.12426","n_code_links":1,"syntology":null},{"paper":null,"title":"HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training","date":"2025-07-15","arxiv_id":"2507.10920","n_code_links":0,"syntology":null},{"paper":null,"title":"Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning","date":"2025-07-14","arxiv_id":"2507.10348","n_code_links":0,"syntology":null},{"paper":null,"title":"SFedKD: Sequential Federated Learning with Discrepancy-Aware Multi-Teacher Knowledge Distillation","date":"2025-07-11","arxiv_id":"2507.08508","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Collaborative Fairness in Federated Learning Under Imbalanced Covariate Shift","date":"2025-07-11","arxiv_id":"2507.08617","n_code_links":0,"syntology":null},{"paper":null,"title":"Continual Self-Supervised Learning with Masked Autoencoders in Remote Sensing","date":"2025-06-26","arxiv_id":"2506.21312","n_code_links":0,"syntology":null},{"paper":null,"title":"Distilling Normalizing Flows","date":"2025-06-26","arxiv_id":"2506.21003","n_code_links":0,"syntology":null},{"paper":"/paper/g-2-d-boosting-multimodal-learning-with","title":"G$^{2}$D: Boosting Multimodal Learning with Gradient-Guided Distillation","date":"2025-06-26","arxiv_id":"2506.21514","n_code_links":1,"syntology":null},{"paper":null,"title":"Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation","date":"2025-06-25","arxiv_id":"2506.20688","n_code_links":0,"syntology":null},{"paper":null,"title":"Client Clustering Meets Knowledge Sharing: Enhancing Privacy and Robustness in Personalized Peer-to-Peer Learning","date":"2025-06-25","arxiv_id":"2506.20413","n_code_links":0,"syntology":null},{"paper":"/paper/fedbkd-distilled-federated-learning-to","title":"FedBKD: Distilled Federated Learning to Embrace Gerneralization and Personalization on Non-IID Data","date":"2025-06-25","arxiv_id":"2506.20245","n_code_links":1,"syntology":null},{"paper":"/paper/tackling-data-heterogeneity-in-federated-2","title":"Tackling Data Heterogeneity in Federated Learning through Knowledge Distillation with Inequitable Aggregation","date":"2025-06-25","arxiv_id":"2506.20431","n_code_links":1,"syntology":null},{"paper":null,"title":"Towards Scalable and Generalizable Earth Observation Data Mining via Foundation Model Composition","date":"2025-06-25","arxiv_id":"2506.20174","n_code_links":0,"syntology":null},{"paper":null,"title":"Distillation-Enabled Knowledge Alignment for Generative Semantic Communications in AIGC Provisioning Tasks","date":"2025-06-24","arxiv_id":"2506.19893","n_code_links":0,"syntology":null},{"paper":null,"title":"Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy","date":"2025-06-24","arxiv_id":"2506.19486","n_code_links":0,"syntology":null},{"paper":null,"title":"PicoSAM2: Low-Latency Segmentation In-Sensor for Edge Vision Applications","date":"2025-06-23","arxiv_id":"2506.18807","n_code_links":0,"syntology":null},{"paper":"/paper/multimodal-fusion-slam-with-fourier-attention","title":"Multimodal Fusion SLAM with Fourier Attention","date":"2025-06-22","arxiv_id":"2506.18204","n_code_links":1,"syntology":null},{"paper":null,"title":"Fine-grained Image Retrieval via Dual-Vision Adaptation","date":"2025-06-19","arxiv_id":"2506.16273","n_code_links":0,"syntology":null},{"paper":null,"title":"Factorized RVQ-GAN For Disentangled Speech Tokenization","date":"2025-06-18","arxiv_id":"2506.15456","n_code_links":0,"syntology":null},{"paper":null,"title":"Knowledge Distillation Framework for Accelerating High-Accuracy Neural Network-Based Molecular Dynamics Simulations","date":"2025-06-18","arxiv_id":"2506.15337","n_code_links":0,"syntology":null},{"paper":null,"title":"AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes","date":"2025-06-17","arxiv_id":"2506.14728","n_code_links":0,"syntology":null},{"paper":"/paper/kdmos-knowledge-distillation-for-motion","title":"KDMOS:Knowledge Distillation for Motion Segmentation","date":"2025-06-17","arxiv_id":"2506.14130","n_code_links":1,"syntology":null},{"paper":null,"title":"Model compression using knowledge distillation with integrated gradients","date":"2025-06-17","arxiv_id":"2506.14440","n_code_links":0,"syntology":null},{"paper":null,"title":"A Technical Study into Small Reasoning Language Models","date":"2025-06-16","arxiv_id":"2506.13404","n_code_links":0,"syntology":null},{"paper":null,"title":"HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs","date":"2025-06-16","arxiv_id":"2506.13038","n_code_links":0,"syntology":null},{"paper":null,"title":"Lightweight Task-Oriented Semantic Communication Empowered by Large-Scale AI Models","date":"2025-06-16","arxiv_id":"2506.13243","n_code_links":0,"syntology":null},{"paper":"/paper/seqpe-transformer-with-sequential-position","title":"SeqPE: Transformer with Sequential Position Encoding","date":"2025-06-16","arxiv_id":"2506.13277","n_code_links":1,"syntology":{"ran":4,"of":12,"unverified":8,"pointer_only":12}},{"paper":null,"title":"Ground Reaction Force Estimation via Time-aware Knowledge Distillation","date":"2025-06-12","arxiv_id":"2506.10265","n_code_links":0,"syntology":null},{"paper":null,"title":"A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning","date":"2025-06-11","arxiv_id":"2506.09429","n_code_links":0,"syntology":null},{"paper":"/paper/2506-08717","title":"Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition","date":"2025-06-10","arxiv_id":"2506.08717","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":3039},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":311},{"task":"/task/model-compression","name":"Model Compression","papers":253},{"task":"/task/object-detection","name":"Object Detection","papers":208},{"task":"/task/object-detection-1","name":"object-detection","papers":205},{"task":"/task/image-classification","name":"Image Classification","papers":202},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":199},{"task":"/task/image-classification","name":"image-classification","papers":179},{"task":"/task/language-modelling","name":"Language Modelling","papers":168},{"task":"/task/language-modeling","name":"Language Modeling","papers":133},{"task":"/task/federated-learning","name":"Federated Learning","papers":123},{"task":"/task/representation-learning","name":"Representation Learning","papers":112},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":111},{"task":"/task/translation","name":"Translation","papers":108},{"task":"/task/segmentation","name":"Segmentation","papers":105},{"task":"/task/object","name":"Object","papers":99},{"task":"/task/incremental-learning","name":"Incremental Learning","papers":98},{"task":"/task/quantization","name":"Quantization","papers":97},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":96},{"task":"/task/continual-learning","name":"Continual Learning","papers":93}],"tasks_shown":20,"n_tasks":1024,"usage_by_year":[{"year":"2015","papers":1},{"year":"2016","papers":4},{"year":"2017","papers":8},{"year":"2018","papers":33},{"year":"2019","papers":130},{"year":"2020","papers":271},{"year":"2021","papers":423},{"year":"2022","papers":506},{"year":"2023","papers":637},{"year":"2024","papers":739},{"year":"2025","papers":319}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/knowledge-distillation"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}