{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-fusion-via-teacher-student-network","title":"Multimodal Fusion via Teacher-Student Network for Indoor Action Recognition","arxiv_id":null,"date":"2021-05-18","proceeding":"Association for the Advancement of Artificial Intelligence (AAAI) 2021 5","authors":["Bruce X.B. Yu","Yan Liu","Keith C.C. Chan"],"abstract":"Indoor action recognition plays an important role in modern\r\nsociety, such as intelligent healthcare in large mobile cabin\r\nhospitals. With the wide usage of depth sensors like Kinect,\r\nmultimodal information including skeleton and RGB modalities\r\nbrings a promising way to improve the performance.\r\nHowever, existing methods are either focusing on a single\r\ndata modality or failed to take the advantage of multiple data\r\nmodalities. In this paper, we propose a Teacher-Student Multimodal\r\nFusion (TSMF) model that fuses the skeleton and\r\nRGB modalities at the model level for indoor action recognition.\r\nIn our TSMF, we utilize a teacher network to transfer\r\nthe structural knowledge of the skeleton modality to a\r\nstudent network for the RGB modality. With extensive experiments\r\non two benchmarking datasets: NTU RGB+D and\r\nPKU-MMD, results show that the proposed TSMF consistently\r\nperforms better than state-of-the-art single modal and\r\nmultimodal methods. It also indicates that our TSMF could\r\nnot only improve the accuracy of the student network but also\r\nsignificantly improve the ensemble accuracy.","url_abs":"https://ojs.aaai.org/index.php/AAAI/article/view/16430","url_pdf":"https://ojs.aaai.org/index.php/AAAI/article/view/16430/16237","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-fusion-via-teacher-student-network","repo_url":"https://github.com/bruceyo/TSMF","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ntu-rgbd","task":"Action Recognition","dataset":"NTU RGB+D","model":"TSMF (RGB + Pose)","rank_in_archive_order":18,"of":28,"metrics":{"Accuracy (CS)":"92.5","Accuracy (CV)":"97.4"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-pku-mmd","task":"Action Recognition In Videos","dataset":"PKU-MMD","model":"TSMF","rank_in_archive_order":4,"of":5,"metrics":{"X-Sub":"95.8","X-View":"97.8"},"uses_additional_data":true}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}