{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sharp-segmentation-of-hands-and-arms-by-range","title":"SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition","arxiv_id":"2408.10037","date":"2024-08-19","proceeding":null,"authors":["Wiktor Mucha","Michael Wray","Martin Kampel"],"abstract":"Hand pose represents key information for action recognition in the egocentric perspective, where the user is interacting with objects. We propose to improve egocentric 3D hand pose estimation based on RGB frames only by using pseudo-depth images. Incorporating state-of-the-art single RGB image depth estimation techniques, we generate pseudo-depth representations of the frames and use distance knowledge to segment irrelevant parts of the scene. The resulting depth maps are then used as segmentation masks for the RGB frames. Experimental results on H2O Dataset confirm the high accuracy of the estimated pose with our method in an action recognition task. The 3D hand pose, together with information from object detection, is processed by a transformer-based action recognition network, resulting in an accuracy of 91.73%, outperforming all state-of-the-art methods. Estimations of 3D hand pose result in competitive performance with existing methods with a mean pose error of 28.66 mm. This method opens up new possibilities for employing distance information in egocentric 3D hand pose estimation without relying on depth sensors.","url_abs":"https://arxiv.org/abs/2408.10037v1","url_pdf":"https://arxiv.org/pdf/2408.10037v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sharp-segmentation-of-hands-and-arms-by-range","repo_url":"https://github.com/wiktormucha/SHARP","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-hand-pose-estimation","task_name":"3D Hand Pose Estimation"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"hand-pose-estimation","task_name":"Hand Pose Estimation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-on-h2o-2-hands-and-objects","task":"Action Recognition","dataset":"H2O  (2 Hands and Objects)","model":"SHARP","rank_in_archive_order":2,"of":11,"metrics":{"Actions Top-1":"91.73","Hand Pose":"3D","Object Label":"Yes","Object Pose":"2D","RGB":"No"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-h2o-2","task":"Skeleton Based Action Recognition","dataset":"H2O  (2 Hands and Objects)","model":"SHARP","rank_in_archive_order":2,"of":4,"metrics":{"Accuracy":"91.73"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}