{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/variational-information-distillation-for","title":"Variational Information Distillation for Knowledge Transfer","arxiv_id":"1904.05835","date":"2019-04-11","proceeding":"CVPR 2019 6","authors":["Sungsoo Ahn","Shell Xu Hu","Andreas Damianou","Neil D. Lawrence","Zhenwen Dai"],"abstract":"Transferring knowledge from a teacher neural network pretrained on the same\nor a similar task to a student neural network can significantly improve the\nperformance of the student neural network. Existing knowledge transfer\napproaches match the activations or the corresponding hand-crafted features of\nthe teacher and the student networks. We propose an information-theoretic\nframework for knowledge transfer which formulates knowledge transfer as\nmaximizing the mutual information between the teacher and the student networks.\nWe compare our method with existing knowledge transfer methods on both\nknowledge distillation and transfer learning tasks and show that our method\nconsistently outperforms existing methods. We further demonstrate the strength\nof our method on knowledge transfer across heterogeneous network architectures\nby transferring knowledge from a convolutional neural network (CNN) to a\nmulti-layer perceptron (MLP) on CIFAR-10. The resulting MLP significantly\noutperforms the-state-of-the-art methods and it achieves similar performance to\nthe CNN with a single convolutional layer.","url_abs":"http://arxiv.org/abs/1904.05835v1","url_pdf":"http://arxiv.org/pdf/1904.05835v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"variational-information-distillation-for","repo_url":"https://github.com/amzn/xfer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mxnet","reach":{"status":"unanswered"}},{"paper_slug":"variational-information-distillation-for","repo_url":"https://github.com/yoshitomo-matsubara/torchdistill","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.05835","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}