{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-relational-modeling-for-action","title":"Relational Network for Skeleton-Based Action Recognition","arxiv_id":"1805.02556","date":"2018-05-07","proceeding":null,"authors":["Wu Zheng","Lin Li","Zhao-Xiang Zhang","Yan Huang","Liang Wang"],"abstract":"With the fast development of effective and low-cost human skeleton capture\nsystems, skeleton-based action recognition has attracted much attention\nrecently. Most existing methods use Convolutional Neural Network (CNN) and\nRecurrent Neural Network (RNN) to extract spatio-temporal information embedded\nin the skeleton sequences for action recognition. However, these approaches are\nlimited in the ability of relational modeling in a single skeleton, due to the\nloss of important structural information when converting the raw skeleton data\nto adapt to the input format of CNN or RNN. In this paper, we propose an\nAttentional Recurrent Relational Network-LSTM (ARRN-LSTM) to simultaneously\nmodel spatial configurations and temporal dynamics in skeletons for action\nrecognition. We introduce the Recurrent Relational Network to learn the spatial\nfeatures in a single skeleton, followed by a multi-layer LSTM to learn the\ntemporal features in the skeleton sequences. Between the two modules, we design\nan adaptive attentional module to focus attention on the most discriminative\nparts in the single skeleton. To exploit the complementarity from different\ngeometries in the skeleton for sufficient relational modeling, we design a\ntwo-stream architecture to learn the structural features among joints and lines\nsimultaneously. Extensive experiments are conducted on several popular skeleton\ndatasets and the results show that the proposed approach achieves better\nresults than most mainstream methods.","url_abs":"http://arxiv.org/abs/1805.02556v4","url_pdf":"http://arxiv.org/pdf/1805.02556v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"ARRN-LSTM","rank_in_archive_order":111,"of":135,"metrics":{"Accuracy (CS)":"80.7","Accuracy (CV)":"88.8"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}