{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unik-a-unified-framework-for-real-world","title":"UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition","arxiv_id":"2107.08580","date":"2021-07-19","proceeding":null,"authors":["Di Yang","Yaohui Wang","Antitza Dantcheva","Lorenzo Garattoni","Gianpiero Francesca","Francois Bremond"],"abstract":"Action recognition based on skeleton data has recently witnessed increasing attention and progress. State-of-the-art approaches adopting Graph Convolutional networks (GCNs) can effectively extract features on human skeletons relying on the pre-defined human topology. Despite associated progress, GCN-based methods have difficulties to generalize across domains, especially with different human topological structures. In this context, we introduce UNIK, a novel skeleton-based action recognition method that is not only effective to learn spatio-temporal features on human skeleton sequences but also able to generalize across datasets. This is achieved by learning an optimal dependency matrix from the uniform distribution based on a multi-head attention mechanism. Subsequently, to study the cross-domain generalizability of skeleton-based action recognition in real-world videos, we re-evaluate state-of-the-art approaches as well as the proposed UNIK in light of a novel Posetics dataset. This dataset is created from Kinetics-400 videos by estimating, refining and filtering poses. We provide an analysis on how much performance improves on smaller benchmark datasets after pre-training on Posetics for the action classification task. Experimental results show that the proposed UNIK, with pre-training on Posetics, generalizes well and outperforms state-of-the-art when transferred onto four target action classification datasets: Toyota Smarthome, Penn Action, NTU-RGB+D 60 and NTU-RGB+D 120.","url_abs":"https://arxiv.org/abs/2107.08580v1","url_pdf":"https://arxiv.org/pdf/2107.08580v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unik-a-unified-framework-for-real-world","repo_url":"https://github.com/YangDi666/UNIK","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"}],"methods":[{"method_slug":"graph-convolutional-networks","method_name":"Graph Convolutional Networks"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-toyota-smarthome","task":"Action Classification","dataset":"Toyota Smarthome dataset","model":"UNIK","rank_in_archive_order":4,"of":13,"metrics":{"CS":"64.3","CV1":"36.1","CV2":"65.0"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-upenn","task":"Skeleton Based Action Recognition","dataset":"UPenn Action","model":"UNIK","rank_in_archive_order":1,"of":3,"metrics":{"Accuracy":"97.9"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2107.08580","atlas_url":"https://app.syntology.ai/?focus=2107.08580","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}