{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-action-recognition-with-1","title":"Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning","arxiv_id":"1805.02335","date":"2018-05-07","proceeding":"ECCV 2018 9","authors":["Chenyang Si","Ya Jing","Wei Wang","Liang Wang","Tieniu Tan"],"abstract":"Skeleton-based action recognition has made great progress recently, but many\nproblems still remain unsolved. For example, most of the previous methods model\nthe representations of skeleton sequences without abundant spatial structure\ninformation and detailed temporal dynamics features. In this paper, we propose\na novel model with spatial reasoning and temporal stack learning (SR-TSL) for\nskeleton based action recognition, which consists of a spatial reasoning\nnetwork (SRN) and a temporal stack learning network (TSLN). The SRN can capture\nthe high-level spatial structural information within each frame by a residual\ngraph neural network, while the TSLN can model the detailed temporal dynamics\nof skeleton sequences by a composition of multiple skip-clip LSTMs. During\ntraining, we propose a clip-based incremental loss to optimize the model. We\nperform extensive experiments on the SYSU 3D Human-Object Interaction dataset\nand NTU RGB+D dataset and verify the effectiveness of each network of our\nmodel. The comparison results illustrate that our approach achieves much better\nresults than state-of-the-art methods.","url_abs":"http://arxiv.org/abs/1805.02335v2","url_pdf":"http://arxiv.org/pdf/1805.02335v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"graph-neural-network","task_name":"Graph Neural Network"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"spatial-reasoning","task_name":"Spatial Reasoning"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"SR-TSL","rank_in_archive_order":97,"of":135,"metrics":{"Accuracy (CS)":"84.8","Accuracy (CV)":"92.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.02335","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}