{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ntu-rgbd-a-large-scale-dataset-for-3d-human","title":"NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis","arxiv_id":"1604.02808","date":"2016-04-11","proceeding":"CVPR 2016 6","authors":["Amir Shahroudy","Jun Liu","Tian-Tsong Ng","Gang Wang"],"abstract":"Recent approaches in depth-based human activity analysis achieved outstanding\nperformance and proved the effectiveness of 3D representation for\nclassification of action classes. Currently available depth-based and\nRGB+D-based action recognition benchmarks have a number of limitations,\nincluding the lack of training samples, distinct class labels, camera views and\nvariety of subjects. In this paper we introduce a large-scale dataset for RGB+D\nhuman action recognition with more than 56 thousand video samples and 4 million\nframes, collected from 40 distinct subjects. Our dataset contains 60 different\naction classes including daily, mutual, and health-related actions. In\naddition, we propose a new recurrent neural network structure to model the\nlong-term temporal correlation of the features for each body part, and utilize\nthem for better action classification. Experimental results show the advantages\nof applying deep learning methods over state-of-the-art hand-crafted features\non the suggested cross-subject and cross-view evaluation criteria for our\ndataset. The introduction of this large scale dataset will enable the community\nto apply, develop and adapt various data-hungry learning techniques for the\ntask of depth-based and RGB+D-based human activity analysis.","url_abs":"http://arxiv.org/abs/1604.02808v1","url_pdf":"http://arxiv.org/pdf/1604.02808v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ntu-rgbd-a-large-scale-dataset-for-3d-human","repo_url":"https://github.com/shahroudy/NTURGB-D","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"ntu-rgbd-a-large-scale-dataset-for-3d-human","repo_url":"https://github.com/nntanaka/Fourier-Analysis-for-Skeleton-based-Action-Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-human-action-recognition","task_name":"3D Action Recognition"},{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"}],"methods":[],"datasets_introduced":[{"slug":"ntu-rgb-d","name":"NTU RGB+D","full_name":""},{"slug":"ntu-rgb-d-2d","name":"NTU RGB+D 2D","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-cad-120","task":"Skeleton Based Action Recognition","dataset":"CAD-120","model":"P-LSTM (5-shot)","rank_in_archive_order":8,"of":8,"metrics":{"Accuracy":"68.1%"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"Part-aware LSTM","rank_in_archive_order":128,"of":135,"metrics":{"Accuracy (CS)":"62.93","Accuracy (CV)":"70.27"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"Deep LSTM","rank_in_archive_order":130,"of":135,"metrics":{"Accuracy (CS)":"60.7","Accuracy (CV)":"67.3"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"Part-Aware LSTM","rank_in_archive_order":83,"of":83,"metrics":{"Accuracy (Cross-Setup)":"26.3%","Accuracy (Cross-Subject)":"25.5%"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-varying","task":"Skeleton Based Action Recognition","dataset":"Varying-view RGB-D Action-Skeleton","model":"P-LSTM","rank_in_archive_order":4,"of":7,"metrics":{"Accuracy (AV I)":"33%","Accuracy (AV II)":"50%","Accuracy (CS)":"60%","Accuracy (CV I)":"13%","Accuracy (CV II)":"33%"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-varying","task":"Skeleton Based Action Recognition","dataset":"Varying-view RGB-D Action-Skeleton","model":"LSTM","rank_in_archive_order":7,"of":7,"metrics":{"Accuracy (AV I)":"31%","Accuracy (AV II)":"68%","Accuracy (CS)":"56%","Accuracy (CV I)":"16%","Accuracy (CV II)":"31%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1604.02808","atlas_url":"https://app.syntology.ai/?focus=1604.02808","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}