{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-attention-enhanced-graph-convolutional","title":"An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition","arxiv_id":"1902.09130","date":"2019-02-25","proceeding":"CVPR 2019 6","authors":["Chenyang Si","Wentao Chen","Wei Wang","Liang Wang","Tieniu Tan"],"abstract":"Skeleton-based action recognition is an important task that requires the\nadequate understanding of movement characteristics of a human action from the\ngiven skeleton sequence. Recent studies have shown that exploring spatial and\ntemporal features of the skeleton sequence is vital for this task.\nNevertheless, how to effectively extract discriminative spatial and temporal\nfeatures is still a challenging problem. In this paper, we propose a novel\nAttention Enhanced Graph Convolutional LSTM Network (AGC-LSTM) for human action\nrecognition from skeleton data. The proposed AGC-LSTM can not only capture\ndiscriminative features in spatial configuration and temporal dynamics but also\nexplore the co-occurrence relationship between spatial and temporal domains. We\nalso present a temporal hierarchical architecture to increases temporal\nreceptive fields of the top AGC-LSTM layer, which boosts the ability to learn\nthe high-level semantic representation and significantly reduces the\ncomputation cost. Furthermore, to select discriminative spatial information,\nthe attention mechanism is employed to enhance information of key joints in\neach AGC-LSTM layer. Experimental results on two datasets are provided: NTU\nRGB+D dataset and Northwestern-UCLA dataset. The comparison results demonstrate\nthe effectiveness of our approach and show that our approach outperforms the\nstate-of-the-art methods on both datasets.","url_abs":"http://arxiv.org/abs/1902.09130v2","url_pdf":"http://arxiv.org/pdf/1902.09130v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"AGC-LSTM (Joint&Part)","rank_in_archive_order":64,"of":135,"metrics":{"Accuracy (CS)":"89.2","Accuracy (CV)":"95.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.09130","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}