{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-end-to-end-spatio-temporal-attention-model","title":"An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data","arxiv_id":"1611.06067","date":"2016-11-18","proceeding":null,"authors":["Sijie Song","Cuiling Lan","Junliang Xing","Wen-Jun Zeng","Jiaying Liu"],"abstract":"Human action recognition is an important task in computer vision. Extracting\ndiscriminative spatial and temporal features to model the spatial and temporal\nevolutions of different actions plays a key role in accomplishing this task. In\nthis work, we propose an end-to-end spatial and temporal attention model for\nhuman action recognition from skeleton data. We build our model on top of the\nRecurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which\nlearns to selectively focus on discriminative joints of skeleton within each\nframe of the inputs and pays different levels of attention to the outputs of\ndifferent frames. Furthermore, to ensure effective training of the network, we\npropose a regularized cross-entropy loss to drive the model learning process\nand develop a joint training strategy accordingly. Experimental results\ndemonstrate the effectiveness of the proposed model,both on the small human\naction recognition data set of SBU and the currently largest NTU dataset.","url_abs":"http://arxiv.org/abs/1611.06067v1","url_pdf":"http://arxiv.org/pdf/1611.06067v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"sta-lstm","method_name":"STA-LSTM"},{"method_slug":"spatial-temporal-attention","method_name":"Spatial & Temporal Attention"}],"datasets_introduced":[],"methods_introduced":[{"slug":"sta-lstm","name":"STA-LSTM","full_name":"Spatio-Temporal Attention LSTM"},{"slug":"spatial-temporal-attention","name":"Spatial & Temporal Attention","full_name":"Spatial & Temporal Attention"}],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"STA-LSTM","rank_in_archive_order":124,"of":135,"metrics":{"Accuracy (CS)":"73.4","Accuracy (CV)":"81.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.06067","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}