{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/jointly-learning-heterogeneous-features-for-1","title":"Jointly learning heterogeneous features for rgb-d activity recognition","arxiv_id":null,"date":"2016-12-15","proceeding":"IEEE Transactions on Pattern Analysis and Machine Intelligence ( Volume: 39 , Issue: 11 , Nov. 1 2017 ) 2016 12","authors":["Jian-Fang Hu","Wei-Shi Zheng","Jian-Huang Lai","Jian-Guo Zhang"],"abstract":"In this paper, we focus on heterogeneous features learning for RGB-D activity recognition. We find that features from different channels (RGB, depth) could share some similar hidden structures, and then propose a joint learning model to simultaneously explore the shared and feature-specific components as an instance of heterogeneous multi-task learning. The proposed model formed in a unified framework is capable of: 1) jointly mining a set of subspaces with the same dimensionality to exploit latent shared features across different feature channels, 2) meanwhile, quantifying the shared and feature-specific components of features in the subspaces, and 3) transferring feature-specific intermediate transforms (i-transforms) for learning fusion of heterogeneous features across datasets. To efficiently train the joint model, a three-step iterative optimization algorithm is proposed, followed by a simple inference model. Extensive experimental results on four activity datasets have demonstrated the efficacy of the proposed method. Anew RGB-D activity dataset focusing on human-object interaction is further contributed, which presents more challenges for RGB-D activity benchmarking.","url_abs":"https://doi.org/10.1109/TPAMI.2016.2640292","url_pdf":"https://www.cv-foundation.org/openaccess/content_cvpr_2015/papers/Hu_Jointly_Learning_Heterogeneous_2015_CVPR_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"Dynamic Skeletons","rank_in_archive_order":131,"of":135,"metrics":{"Accuracy (CS)":"60.2","Accuracy (CV)":"65.2"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"Dynamic Skeletons","rank_in_archive_order":81,"of":83,"metrics":{"Accuracy (Cross-Setup)":"54.7%","Accuracy (Cross-Subject)":"50.8%"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-sysu-3d","task":"Skeleton Based Action Recognition","dataset":"SYSU 3D","model":"Dynamic Skeletons","rank_in_archive_order":8,"of":9,"metrics":{"Accuracy":"75.5%"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}