{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-unconstrained-joint-hand-object","title":"Towards unconstrained joint hand-object reconstruction from RGB videos","arxiv_id":"2108.07044","date":"2021-08-16","proceeding":null,"authors":["Yana Hasson","Gül Varol","Ivan Laptev","Cordelia Schmid"],"abstract":"Our work aims to obtain 3D reconstruction of hands and manipulated objects from monocular videos. Reconstructing hand-object manipulations holds a great potential for robotics and learning from human demonstrations. The supervised learning approach to this problem, however, requires 3D supervision and remains limited to constrained laboratory settings and simulators for which 3D ground truth is available. In this paper we first propose a learning-free fitting approach for hand-object reconstruction which can seamlessly handle two-hand object interactions. Our method relies on cues obtained with common methods for object detection, hand pose estimation and instance segmentation. We quantitatively evaluate our approach and show that it can be applied to datasets with varying levels of difficulty for which training data is unavailable.","url_abs":"https://arxiv.org/abs/2108.07044v2","url_pdf":"https://arxiv.org/pdf/2108.07044v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-unconstrained-joint-hand-object","repo_url":"https://github.com/hassony2/homan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-hand-pose-estimation","task_name":"3D Hand Pose Estimation"},{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"},{"task_slug":"hand-pose-estimation","task_name":"Hand Pose Estimation"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-reconstruction","task_name":"Object Reconstruction"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"hand-object-pose","task_name":"hand-object pose"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-hand-pose-estimation-on-ho-3d","task":"3D Hand Pose Estimation","dataset":"HO-3D v2","model":"HOR","rank_in_archive_order":22,"of":24,"metrics":{"PA-MPJPE (mm)":"12.0"},"uses_additional_data":false},{"leaderboard":"/sota/hand-object-pose-on-dexycb","task":"hand-object pose","dataset":"DexYCB","model":"UHO","rank_in_archive_order":8,"of":9,"metrics":{"ADD-S":"-","Average MPJPE (mm)":"18.8","MCE":"52.5","OCE":"-","Procrustes-Aligned MPJPE":"-"},"uses_additional_data":false},{"leaderboard":"/sota/hand-object-pose-on-ho-3d","task":"hand-object pose","dataset":"HO-3D v2","model":"HOR","rank_in_archive_order":5,"of":9,"metrics":{"ADD-S":"40.0","Average MPJPE (mm)":"-","OME":"80.0","PA-MPJPE":"12.0","ST-MPJPE":"26.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2108.07044","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}