{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-new-representation-of-skeleton-sequences","title":"A New Representation of Skeleton Sequences for 3D Action Recognition","arxiv_id":"1703.03492","date":"2017-03-09","proceeding":"CVPR 2017 7","authors":["Qiuhong Ke","Mohammed Bennamoun","Senjian An","Ferdous Sohel","Farid Boussaid"],"abstract":"This paper presents a new method for 3D action recognition with skeleton\nsequences (i.e., 3D trajectories of human skeleton joints). The proposed method\nfirst transforms each skeleton sequence into three clips each consisting of\nseveral frames for spatial temporal feature learning using deep neural\nnetworks. Each clip is generated from one channel of the cylindrical\ncoordinates of the skeleton sequence. Each frame of the generated clips\nrepresents the temporal information of the entire skeleton sequence, and\nincorporates one particular spatial relationship between the joints. The entire\nclips include multiple frames with different spatial relationships, which\nprovide useful spatial structural information of the human skeleton. We propose\nto use deep convolutional neural networks to learn long-term temporal\ninformation of the skeleton sequence from the frames of the generated clips,\nand then use a Multi-Task Learning Network (MTLN) to jointly process all frames\nof the generated clips in parallel to incorporate spatial structural\ninformation for action recognition. Experimental results clearly show the\neffectiveness of the proposed new representation and feature learning method\nfor 3D action recognition.","url_abs":"http://arxiv.org/abs/1703.03492v3","url_pdf":"http://arxiv.org/pdf/1703.03492v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-action-recognition","task_name":"3D Action Recognition"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"Clips+CNN+MTLN","rank_in_archive_order":115,"of":135,"metrics":{"Accuracy (CS)":"79.6","Accuracy (CV)":"84.8"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"Multi-Task Learning Network","rank_in_archive_order":77,"of":83,"metrics":{"Accuracy (Cross-Setup)":"57.9%","Accuracy (Cross-Subject)":"58.4%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.03492","atlas_url":"https://app.syntology.ai/?focus=1703.03492","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}