{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-human-action-recognition-with","title":"Skeleton-Based Human Action Recognition with Global Context-Aware Attention LSTM Networks","arxiv_id":"1707.05740","date":"2017-07-18","proceeding":null,"authors":["Jun Liu","Gang Wang","Ling-Yu Duan","Kamila Abdiyeva","Alex C. Kot"],"abstract":"Human action recognition in 3D skeleton sequences has attracted a lot of\nresearch attention. Recently, Long Short-Term Memory (LSTM) networks have shown\npromising performance in this task due to their strengths in modeling the\ndependencies and dynamics in sequential data. As not all skeletal joints are\ninformative for action recognition, and the irrelevant joints often bring noise\nwhich can degrade the performance, we need to pay more attention to the\ninformative ones. However, the original LSTM network does not have explicit\nattention ability. In this paper, we propose a new class of LSTM network,\nGlobal Context-Aware Attention LSTM (GCA-LSTM), for skeleton based action\nrecognition. This network is capable of selectively focusing on the informative\njoints in each frame of each skeleton sequence by using a global context memory\ncell. To further improve the attention capability of our network, we also\nintroduce a recurrent attention mechanism, with which the attention performance\nof the network can be enhanced progressively. Moreover, we propose a stepwise\ntraining scheme in order to train our network effectively. Our approach\nachieves state-of-the-art performance on five challenging benchmark datasets\nfor skeleton based action recognition.","url_abs":"http://arxiv.org/abs/1707.05740v5","url_pdf":"http://arxiv.org/pdf/1707.05740v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"Two-Stream Attention LSTM","rank_in_archive_order":74,"of":83,"metrics":{"Accuracy (Cross-Setup)":"63.3%","Accuracy (Cross-Subject)":"61.2%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.05740","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}