{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-multimodal-feature-analysis-for-action","title":"Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos","arxiv_id":"1603.07120","date":"2016-03-23","proceeding":null,"authors":["Amir Shahroudy","Tian-Tsong Ng","Yihong Gong","Gang Wang"],"abstract":"Single modality action recognition on RGB or depth sequences has been\nextensively explored recently. It is generally accepted that each of these two\nmodalities has different strengths and limitations for the task of action\nrecognition. Therefore, analysis of the RGB+D videos can help us to better\nstudy the complementary properties of these two types of modalities and achieve\nhigher levels of performance. In this paper, we propose a new deep autoencoder\nbased shared-specific feature factorization network to separate input\nmultimodal signals into a hierarchy of components. Further, based on the\nstructure of the features, a structured sparsity learning machine is proposed\nwhich utilizes mixed norms to apply regularization within components and group\nselection between them for better classification performance. Our experimental\nresults show the effectiveness of our cross-modality feature analysis framework\nby achieving state-of-the-art accuracy for action classification on five\nchallenging benchmark datasets.","url_abs":"http://arxiv.org/abs/1603.07120v2","url_pdf":"http://arxiv.org/pdf/1603.07120v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"multimodal-activity-recognition","task_name":"Multimodal Activity Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ntu-rgbd","task":"Action Recognition","dataset":"NTU RGB+D","model":"DSSCA-SSLM (RGB only)","rank_in_archive_order":28,"of":28,"metrics":{"Accuracy (CS)":"74.9"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-activity-recognition-on-msr-daily-1","task":"Multimodal Activity Recognition","dataset":"MSR Daily Activity3D dataset","model":"DSSCA-SSLM (RGB+D)","rank_in_archive_order":1,"of":6,"metrics":{"Accuracy":"97.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1603.07120","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}