{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-recognition-based-on-joint-trajectory","title":"Action Recognition Based on Joint Trajectory Maps with Convolutional Neural Networks","arxiv_id":"1612.09401","date":"2016-12-30","proceeding":null,"authors":["Pichao Wang","Wanqing Li","Chuankun Li","Yonghong Hou"],"abstract":"Convolutional Neural Networks (ConvNets) have recently shown promising\nperformance in many computer vision tasks, especially image-based recognition.\nHow to effectively apply ConvNets to sequence-based data is still an open\nproblem. This paper proposes an effective yet simple method to represent\nspatio-temporal information carried in $3D$ skeleton sequences into three $2D$\nimages by encoding the joint trajectories and their dynamics into color\ndistribution in the images, referred to as Joint Trajectory Maps (JTM), and\nadopts ConvNets to learn the discriminative features for human action\nrecognition. Such an image-based representation enables us to fine-tune\nexisting ConvNets models for the classification of skeleton sequences without\ntraining the networks afresh. The three JTMs are generated in three orthogonal\nplanes and provide complimentary information to each other. The final\nrecognition is further improved through multiply score fusion of the three\nJTMs. The proposed method was evaluated on four public benchmark datasets, the\nlarge NTU RGB+D Dataset, MSRC-12 Kinect Gesture Dataset (MSRC-12), G3D Dataset\nand UTD Multimodal Human Action Dataset (UTD-MHAD) and achieved the\nstate-of-the-art results.","url_abs":"http://arxiv.org/abs/1612.09401v1","url_pdf":"http://arxiv.org/pdf/1612.09401v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-gaming","task":"Skeleton Based Action Recognition","dataset":"Gaming 3D (G3D)","model":"CNN","rank_in_archive_order":1,"of":4,"metrics":{"Accuracy":"96.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}