{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skeleton-based-action-recognition-of-people","title":"Skeleton-based Action Recognition of People Handling Objects","arxiv_id":"1901.06882","date":"2019-01-21","proceeding":null,"authors":["Sunoh Kim","Kimin Yun","Jongyoul Park","Jin Young Choi"],"abstract":"In visual surveillance systems, it is necessary to recognize the behavior of\npeople handling objects such as a phone, a cup, or a plastic bag. In this\npaper, to address this problem, we propose a new framework for recognizing\nobject-related human actions by graph convolutional networks using human and\nobject poses. In this framework, we construct skeletal graphs of reliable human\nposes by selectively sampling the informative frames in a video, which include\nhuman joints with high confidence scores obtained in pose estimation. The\nskeletal graphs generated from the sampled frames represent human poses related\nto the object position in both the spatial and temporal domains, and these\ngraphs are used as inputs to the graph convolutional networks. Through\nexperiments over an open benchmark and our own data sets, we verify the\nvalidity of our framework in that our method outperforms the state-of-the-art\nmethod for skeleton-based action recognition.","url_abs":"http://arxiv.org/abs/1901.06882v1","url_pdf":"http://arxiv.org/pdf/1901.06882v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"object","task_name":"Object"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"graph-convolutional-networks","method_name":"Graph Convolutional Networks"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-icvl-4","task":"Action Recognition","dataset":"ICVL-4","model":"OHA-GCN (Two stream; HP + OHP-hands + informative samples)","rank_in_archive_order":1,"of":2,"metrics":{"Accuracy":"91.86%"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ird","task":"Action Recognition","dataset":"IRD","model":"OHA-GCN (Two stream; HP + OHP-hands + informative samples)","rank_in_archive_order":1,"of":2,"metrics":{"Accuracy":"80.11%"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}