{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-action-recognition-using","title":"Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust Gates","arxiv_id":"1706.08276","date":"2017-06-26","proceeding":null,"authors":["Jun Liu","Amir Shahroudy","Dong Xu","Alex C. Kot","Gang Wang"],"abstract":"Skeleton-based human action recognition has attracted a lot of research\nattention during the past few years. Recent works attempted to utilize\nrecurrent neural networks to model the temporal dependencies between the 3D\npositional configurations of human body joints for better analysis of human\nactivities in the skeletal data. The proposed work extends this idea to spatial\ndomain as well as temporal domain to better analyze the hidden sources of\naction-related information within the human skeleton sequences in both of these\ndomains simultaneously. Based on the pictorial structure of Kinect's skeletal\ndata, an effective tree-structure based traversal framework is also proposed.\nIn order to deal with the noise in the skeletal data, a new gating mechanism\nwithin LSTM module is introduced, with which the network can learn the\nreliability of the sequential data and accordingly adjust the effect of the\ninput data on the updating procedure of the long-term context representation\nstored in the unit's memory cell. Moreover, we introduce a novel multi-modal\nfeature fusion strategy within the LSTM unit in this paper. The comprehensive\nexperimental results on seven challenging benchmark datasets for human action\nrecognition demonstrate the effectiveness of the proposed method.","url_abs":"http://arxiv.org/abs/1706.08276v1","url_pdf":"http://arxiv.org/pdf/1706.08276v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"one-shot-3d-action-recognition","task_name":"One-Shot 3D Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/one-shot-3d-action-recognition-on-ntu-rgbd","task":"One-Shot 3D Action Recognition","dataset":"NTU RGB+D 120","model":"Average Pooling","rank_in_archive_order":8,"of":10,"metrics":{"Accuracy":"42.9%"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"Internal Feature Fusion","rank_in_archive_order":79,"of":83,"metrics":{"Accuracy (Cross-Setup)":"60.9%","Accuracy (Cross-Subject)":"58.2%"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-sysu-3d","task":"Skeleton Based Action Recognition","dataset":"SYSU 3D","model":"ST-LSTM (Tree)","rank_in_archive_order":9,"of":9,"metrics":{"Accuracy":"73.4%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.08276","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}