{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/real-time-hand-tracking-under-occlusion-from","title":"Real-time Hand Tracking under Occlusion from an Egocentric RGB-D Sensor","arxiv_id":"1704.02201","date":"2017-04-07","proceeding":"ICCV 2017 10","authors":["Franziska Mueller","Dushyant Mehta","Oleksandr Sotnychenko","Srinath Sridhar","Dan Casas","Christian Theobalt"],"abstract":"We present an approach for real-time, robust and accurate hand pose\nestimation from moving egocentric RGB-D cameras in cluttered real environments.\nExisting methods typically fail for hand-object interactions in cluttered\nscenes imaged from egocentric viewpoints, common for virtual or augmented\nreality applications. Our approach uses two subsequently applied Convolutional\nNeural Networks (CNNs) to localize the hand and regress 3D joint locations.\nHand localization is achieved by using a CNN to estimate the 2D position of the\nhand center in the input, even in the presence of clutter and occlusions. The\nlocalized hand position, together with the corresponding input depth value, is\nused to generate a normalized cropped image that is fed into a second CNN to\nregress relative 3D hand joint locations in real time. For added accuracy,\nrobustness and temporal stability, we refine the pose estimates using a\nkinematic pose tracking energy. To train the CNNs, we introduce a new\nphotorealistic dataset that uses a merged reality approach to capture and\nsynthesize large amounts of annotated data of natural hand interaction in\ncluttered scenes. Through quantitative and qualitative evaluation, we show that\nour method is robust to self-occlusion and occlusions by objects, particularly\nin moving egocentric perspectives.","url_abs":"http://arxiv.org/abs/1704.02201v2","url_pdf":"http://arxiv.org/pdf/1704.02201v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"hand-pose-estimation","task_name":"Hand Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"pose-tracking","task_name":"Pose Tracking"},{"task_slug":null,"task_name":"Position"}],"methods":[],"datasets_introduced":[{"slug":"egodexter","name":"EgoDexter","full_name":"EgoDexter"},{"slug":"synthhands","name":"SynthHands","full_name":"SynthHands"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1704.02201","atlas_url":"https://app.syntology.ai/?focus=1704.02201","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}