{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-multipart-learning-for-action","title":"Multimodal Multipart Learning for Action Recognition in Depth Videos","arxiv_id":"1507.08761","date":"2015-07-31","proceeding":null,"authors":["Amir Shahroudy","Gang Wang","Tian-Tsong Ng","Qingxiong Yang"],"abstract":"The articulated and complex nature of human actions makes the task of action\nrecognition difficult. One approach to handle this complexity is dividing it to\nthe kinetics of body parts and analyzing the actions based on these partial\ndescriptors. We propose a joint sparse regression based learning method which\nutilizes the structured sparsity to model each action as a combination of\nmultimodal features from a sparse set of body parts. To represent dynamics and\nappearance of parts, we employ a heterogeneous set of depth and skeleton based\nfeatures. The proper structure of multimodal multipart features are formulated\ninto the learning framework via the proposed hierarchical mixed norm, to\nregularize the structured features of each part and to apply sparsity between\nthem, in favor of a group feature selection. Our experimental results expose\nthe effectiveness of the proposed learning method in which it outperforms other\nmethods in all three tested datasets while saturating one of them by achieving\nperfect accuracy.","url_abs":"http://arxiv.org/abs/1507.08761v1","url_pdf":"http://arxiv.org/pdf/1507.08761v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"multimodal-activity-recognition","task_name":"Multimodal Activity Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-activity-recognition-on-msr-daily-1","task":"Multimodal Activity Recognition","dataset":"MSR Daily Activity3D dataset","model":"MMMP (Pose+D)","rank_in_archive_order":4,"of":6,"metrics":{"Accuracy":"91.3"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}