{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-action-recognition-with-2","title":"Skeleton-based Action Recognition with Convolutional Neural Networks","arxiv_id":"1704.07595","date":"2017-04-25","proceeding":null,"authors":["Chao Li","Qiaoyong Zhong","Di Xie","ShiLiang Pu"],"abstract":"Current state-of-the-art approaches to skeleton-based action recognition are\nmostly based on recurrent neural networks (RNN). In this paper, we propose a\nnovel convolutional neural networks (CNN) based framework for both action\nclassification and detection. Raw skeleton coordinates as well as skeleton\nmotion are fed directly into CNN for label prediction. A novel skeleton\ntransformer module is designed to rearrange and select important skeleton\njoints automatically. With a simple 7-layer network, we obtain 89.3% accuracy\non validation set of the NTU RGB+D dataset. For action detection in untrimmed\nvideos, we develop a window proposal network to extract temporal segment\nproposals, which are further classified within the same network. On the recent\nPKU-MMD dataset, we achieve 93.7% mAP, surpassing the baseline by a large\nmargin.","url_abs":"http://arxiv.org/abs/1704.07595v1","url_pdf":"http://arxiv.org/pdf/1704.07595v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"skeleton-based-action-recognition-with-2","repo_url":"https://github.com/hikvision-research/skelact","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"CNN+Motion+Trans","rank_in_archive_order":102,"of":135,"metrics":{"Accuracy (CS)":"83.2","Accuracy (CV)":"89.3"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-pku-mmd","task":"Skeleton Based Action Recognition","dataset":"PKU-MMD","model":"Li et al. [[Li et al.2017b]]","rank_in_archive_order":3,"of":4,"metrics":{"mAP@0.50 (CS)":"90.4","mAP@0.50 (CV)":"93.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.07595","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}