{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/amalgamating-knowledge-towards-comprehensive","title":"Amalgamating Knowledge towards Comprehensive Classification","arxiv_id":"1811.02796","date":"2018-11-07","proceeding":null,"authors":["Chengchao Shen","Xinchao Wang","Jie Song","Li Sun","Mingli Song"],"abstract":"With the rapid development of deep learning, there have been an\nunprecedentedly large number of trained deep network models available online.\nReusing such trained models can significantly reduce the cost of training the\nnew models from scratch, if not infeasible at all as the annotations used for\nthe training original networks are often unavailable to public. We propose in\nthis paper to study a new model-reusing task, which we term as \\emph{knowledge\namalgamation}. Given multiple trained teacher networks, each of which\nspecializes in a different classification problem, the goal of knowledge\namalgamation is to learn a lightweight student model capable of handling the\ncomprehensive classification. We assume no other annotations except the outputs\nfrom the teacher models are available, and thus focus on extracting and\namalgamating knowledge from the multiple teachers. To this end, we propose a\npilot two-step strategy to tackle the knowledge amalgamation task, by learning\nfirst the compact feature representations from teachers and then the network\nparameters in a layer-wise manner so as to build the student model. We apply\nthis approach to four public datasets and obtain very encouraging results: even\nwithout any human annotation, the obtained student model is competent to handle\nthe comprehensive classification task and in most cases outperforms the\nteachers in individual sub-tasks.","url_abs":"http://arxiv.org/abs/1811.02796v2","url_pdf":"http://arxiv.org/pdf/1811.02796v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"amalgamating-knowledge-towards-comprehensive","repo_url":"https://github.com/zju-vipa/KamalEngine","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.02796","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}