{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-online-action-prediction-using","title":"Skeleton-Based Online Action Prediction Using Scale Selection Network","arxiv_id":"1902.03084","date":"2019-02-08","proceeding":null,"authors":["Jun Liu","Amir Shahroudy","Gang Wang","Ling-Yu Duan","Alex C. Kot"],"abstract":"Action prediction is to recognize the class label of an ongoing activity when\nonly a part of it is observed. In this paper, we focus on online action\nprediction in streaming 3D skeleton sequences. A dilated convolutional network\nis introduced to model the motion dynamics in temporal dimension via a sliding\nwindow over the temporal axis. Since there are significant temporal scale\nvariations in the observed part of the ongoing action at different time steps,\na novel window scale selection method is proposed to make our network focus on\nthe performed part of the ongoing action and try to suppress the possible\nincoming interference from the previous actions at each step. An activation\nsharing scheme is also proposed to handle the overlapping computations among\nthe adjacent time steps, which enables our framework to run more efficiently.\nMoreover, to enhance the performance of our framework for action prediction\nwith the skeletal input data, a hierarchy of dilated tree convolutions are also\ndesigned to learn the multi-level structured semantic representations over the\nskeleton joints at each frame. Our proposed approach is evaluated on four\nchallenging datasets. The extensive experiments demonstrate the effectiveness\nof our method for skeleton-based online action prediction.","url_abs":"http://arxiv.org/abs/1902.03084v2","url_pdf":"http://arxiv.org/pdf/1902.03084v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"FSNet","rank_in_archive_order":76,"of":83,"metrics":{"Accuracy (Cross-Setup)":"62.4%","Accuracy (Cross-Subject)":"59.9%"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1902.03084","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}