{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dstsa-gcn-advancing-skeleton-based-gesture","title":"DSTSA-GCN: Advancing Skeleton-Based Gesture Recognition with Semantic-Aware Spatio-Temporal Topology Modeling","arxiv_id":"2501.12086","date":"2025-01-21","proceeding":null,"authors":["Hu Cui","Renjing Huang","Ruoyu Zhang","Tessai Hayama"],"abstract":"Graph convolutional networks (GCNs) have emerged as a powerful tool for skeleton-based action and gesture recognition, thanks to their ability to model spatial and temporal dependencies in skeleton data. However, existing GCN-based methods face critical limitations: (1) they lack effective spatio-temporal topology modeling that captures dynamic variations in skeletal motion, and (2) they struggle to model multiscale structural relationships beyond local joint connectivity. To address these issues, we propose a novel framework called Dynamic Spatial-Temporal Semantic Awareness Graph Convolutional Network (DSTSA-GCN). DSTSA-GCN introduces three key modules: Group Channel-wise Graph Convolution (GC-GC), Group Temporal-wise Graph Convolution (GT-GC), and Multi-Scale Temporal Convolution (MS-TCN). GC-GC and GT-GC operate in parallel to independently model channel-specific and frame-specific correlations, enabling robust topology learning that accounts for temporal variations. Additionally, both modules employ a grouping strategy to adaptively capture multiscale structural relationships. Complementing this, MS-TCN enhances temporal modeling through group-wise temporal convolutions with diverse receptive fields. Extensive experiments demonstrate that DSTSA-GCN significantly improves the topology modeling capabilities of GCNs, achieving state-of-the-art performance on benchmark datasets for gesture and action recognition, including SHREC17 Track, DHG-14\\/28, NTU-RGB+D, and NTU-RGB+D-120.","url_abs":"https://arxiv.org/abs/2501.12086v1","url_pdf":"https://arxiv.org/pdf/2501.12086v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dstsa-gcn-advancing-skeleton-based-gesture","repo_url":"https://github.com/HuCui2022/DSTSA-GCN_Gesture","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"gesture-recognition","task_name":"Gesture Recognition"},{"task_slug":"hand-gesture-recognition","task_name":"Hand Gesture Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"}],"methods":[{"method_slug":"gcn","method_name":"GCN"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ntu-rgbd","task":"Action Recognition","dataset":"NTU RGB+D","model":"DSTSA-GCN","rank_in_archive_order":17,"of":28,"metrics":{"Accuracy (CS)":"92.78","Accuracy (CV)":"97.03"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ntu-rgbd-120","task":"Action Recognition","dataset":"NTU RGB+D 120","model":"DSTSA-GCN","rank_in_archive_order":12,"of":21,"metrics":{"Accuracy (Cross-Setup)":"90.97","Accuracy (Cross-Subject)":"89.12"},"uses_additional_data":false},{"leaderboard":"/sota/hand-gesture-recognition-on-dhg-14","task":"Hand Gesture Recognition","dataset":"DHG-14","model":"DSTSA-GCN","rank_in_archive_order":2,"of":13,"metrics":{"Accuracy":"95.04"},"uses_additional_data":false},{"leaderboard":"/sota/hand-gesture-recognition-on-dhg-28","task":"Hand Gesture Recognition","dataset":"DHG-28","model":"DSTSA-GCN","rank_in_archive_order":1,"of":9,"metrics":{"Accuracy":"93.57"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-n-ucla","task":"Skeleton Based Action Recognition","dataset":"N-UCLA","model":"DSTSA-GCN","rank_in_archive_order":11,"of":25,"metrics":{"Accuracy":"96.98"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D","model":"DSTSA-GCN","rank_in_archive_order":30,"of":135,"metrics":{"Accuracy (CS)":"92.78","Accuracy (CV)":"97.03","Ensembled Modalities":"4"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-ntu-rgbd-1","task":"Skeleton Based Action Recognition","dataset":"NTU RGB+D 120","model":"DSTSA-GCN","rank_in_archive_order":22,"of":83,"metrics":{"Accuracy (Cross-Setup)":"90.97","Accuracy (Cross-Subject)":"89.12","Ensembled Modalities":"4"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-shrec","task":"Skeleton Based Action Recognition","dataset":"SHREC 2017 track on 3D Hand Gesture Recognition","model":"DSTSA-GCN","rank_in_archive_order":2,"of":7,"metrics":{"14 gestures accuracy":"97.74","28 gestures accuracy":"95.37"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}