{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/knowledge-distillation-via-route-constrained","title":"Knowledge Distillation via Route Constrained Optimization","arxiv_id":"1904.09149","date":"2019-04-19","proceeding":"ICCV 2019 10","authors":["Xiao Jin","Baoyun Peng","Yi-Chao Wu","Yu Liu","Jiaheng Liu","Ding Liang","Xiaolin Hu"],"abstract":"Distillation-based learning boosts the performance of the miniaturized neural\nnetwork based on the hypothesis that the representation of a teacher model can\nbe used as structured and relatively weak supervision, and thus would be easily\nlearned by a miniaturized model. However, we find that the representation of a\nconverged heavy model is still a strong constraint for training a small student\nmodel, which leads to a high lower bound of congruence loss. In this work,\ninspired by curriculum learning we consider the knowledge distillation from the\nperspective of curriculum learning by routing. Instead of supervising the\nstudent model with a converged teacher model, we supervised it with some anchor\npoints selected from the route in parameter space that the teacher model passed\nby, as we called route constrained optimization (RCO). We experimentally\ndemonstrate this simple operation greatly reduces the lower bound of congruence\nloss for knowledge distillation, hint and mimicking learning. On close-set\nclassification tasks like CIFAR100 and ImageNet, RCO improves knowledge\ndistillation by 2.14% and 1.5% respectively. For the sake of evaluating the\ngeneralization, we also test RCO on the open-set face recognition task\nMegaFace.","url_abs":"http://arxiv.org/abs/1904.09149v1","url_pdf":"http://arxiv.org/pdf/1904.09149v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"knowledge-distillation-via-route-constrained","repo_url":"https://github.com/SforAiDl/KD_Lib","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"}],"methods":[{"method_slug":"knowledge-distillation","method_name":"Knowledge Distillation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.09149","atlas_url":"https://app.syntology.ai/?focus=1904.09149","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}