{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modeling-temporal-dynamics-and-spatial","title":"Modeling Temporal Dynamics and Spatial Configurations of Actions Using Two-Stream Recurrent Neural Networks","arxiv_id":"1704.02581","date":"2017-04-09","proceeding":"CVPR 2017 7","authors":["Hongsong Wang","Liang Wang"],"abstract":"Recently, skeleton based action recognition gains more popularity due to\ncost-effective depth sensors coupled with real-time skeleton estimation\nalgorithms. Traditional approaches based on handcrafted features are limited to\nrepresent the complexity of motion patterns. Recent methods that use Recurrent\nNeural Networks (RNN) to handle raw skeletons only focus on the contextual\ndependency in the temporal domain and neglect the spatial configurations of\narticulated skeletons. In this paper, we propose a novel two-stream RNN\narchitecture to model both temporal dynamics and spatial configurations for\nskeleton based action recognition. We explore two different structures for the\ntemporal stream: stacked RNN and hierarchical RNN. Hierarchical RNN is designed\naccording to human body kinematics. We also propose two effective methods to\nmodel the spatial structure by converting the spatial graph into a sequence of\njoints. To improve generalization of our model, we further exploit 3D\ntransformation based data augmentation techniques including rotation and\nscaling transformation to transform the 3D coordinates of skeletons during\ntraining. Experiments on 3D action recognition benchmark datasets show that our\nmethod brings a considerable improvement for a variety of actions, i.e.,\ngeneric actions, interaction activities and gestures.","url_abs":"http://arxiv.org/abs/1704.02581v2","url_pdf":"http://arxiv.org/pdf/1704.02581v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-action-recognition","task_name":"3D Action Recognition"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"Two-Stream RNN","rank_in_archive_order":126,"of":135,"metrics":{"Accuracy (CS)":"71.3","Accuracy (CV)":"79.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.02581","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}