{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sign-language-recognition-via-deformable-3d","title":"Sign Language Recognition via Deformable 3D Convolutions and Modulated Graph Convolutional Networks","arxiv_id":null,"date":"2023-06-07","proceeding":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2023 6","authors":["Katerina Papadimitriou","Gerasimos Potamianos"],"abstract":"Automatic sign language recognition (SLR) remains challenging, especially when employing RGB video alone (i.e., with no depth or special glove-based input) and under a signer-independent (SI) framework, due to inter-personal signing variation. In this paper, we address SI isolated SLR from RGB video, proposing an innovative deep-learning framework that leverages multi-modal appearanceand skeleton-based information. Specifically, we propose three components for the first time in SLR: (i) a modified version of the ResNet2+1D network to capture signing appearance information, where spatial and temporal convolutions are substituted by their deformable counterparts, accomplishing both prevalent spatial modeling potential and motion-aware modeling adaptability; (ii) a novel spatio-temporal graph convolutional network (ST-GCN) that integrates a GCN variant, involving weight and affinity modulation for modeling diverse correlations between different body joints beyond the physical human skeleton structure, followed by a self-attention layer and a temporal convolution; and (iii) the “PIXIE” 3D human pose and shape regressor to generate 3D joint-rotation parameterization used for ST-GCN graph construction. Both appearance- and skeleton-based streams are ensembled in the proposed system and evaluated on two datasets of isolated signs, one in Turkish and one in Greek. Our system outperforms the state-of-the-art on the second set, yielding 53% relative error rate reduction (2.45% absolute), while it performs on par with the best reported system on the first.","url_abs":"https://ieeexplore.ieee.org/document/10096714","url_pdf":"https://ieeexplore.ieee.org/document/10096714","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"sign-language-recognition","task_name":"Sign Language Recognition"},{"task_slug":"graph-construction","task_name":"graph construction"}],"methods":[{"method_slug":"gcn","method_name":"GCN"},{"method_slug":"slr","method_name":"SLR"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/sign-language-recognition-on-autsl","task":"Sign Language Recognition","dataset":"AUTSL","model":"3D-DCNN + ST-MGCN","rank_in_archive_order":3,"of":9,"metrics":{"Rank-1 Recognition Rate":"0.9842"},"uses_additional_data":false},{"leaderboard":"/sota/sign-language-recognition-on-gsl","task":"Sign Language Recognition","dataset":"GSL","model":"3D-DCNN + ST-MGCN","rank_in_archive_order":1,"of":1,"metrics":{"Rank-1 Recognition Rate":"0.9785"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}