{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-fine-to-coarse-convolutional-neural-network","title":"A Fine-to-Coarse Convolutional Neural Network for 3D Human Action Recognition","arxiv_id":"1805.11790","date":"2018-05-30","proceeding":null,"authors":["Thao Minh Le","Nakamasa Inoue","Koichi Shinoda"],"abstract":"This paper presents a new framework for human action recognition from a 3D\nskeleton sequence. Previous studies do not fully utilize the temporal\nrelationships between video segments in a human action. Some studies\nsuccessfully used very deep Convolutional Neural Network (CNN) models but often\nsuffer from the data insufficiency problem. In this study, we first segment a\nskeleton sequence into distinct temporal segments in order to exploit the\ncorrelations between them. The temporal and spatial features of a skeleton\nsequence are then extracted simultaneously by utilizing a fine-to-coarse (F2C)\nCNN architecture optimized for human skeleton sequences. We evaluate our\nproposed method on NTU RGB+D and SBU Kinect Interaction dataset. It achieves\n79.6% and 84.6% of accuracies on NTU RGB+D with cross-object and cross-view\nprotocol, respectively, which are almost identical with the state-of-the-art\nperformance. In addition, our method significantly improves the accuracy of the\nactions in two-person interactions.","url_abs":"http://arxiv.org/abs/1805.11790v2","url_pdf":"http://arxiv.org/pdf/1805.11790v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-action-recognition","task_name":"3D Action Recognition"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"F2CSkeleton","rank_in_archive_order":116,"of":135,"metrics":{"Accuracy (CS)":"79.6","Accuracy (CV)":"84.6"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}