{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monocular-total-capture-posing-face-body-and","title":"Monocular Total Capture: Posing Face, Body, and Hands in the Wild","arxiv_id":"1812.01598","date":"2018-12-04","proceeding":"CVPR 2019 6","authors":["Donglai Xiang","Hanbyul Joo","Yaser Sheikh"],"abstract":"We present the first method to capture the 3D total motion of a target person\nfrom a monocular view input. Given an image or a monocular video, our method\nreconstructs the motion from body, face, and fingers represented by a 3D\ndeformable mesh model. We use an efficient representation called 3D Part\nOrientation Fields (POFs), to encode the 3D orientations of all body parts in\nthe common 2D image space. POFs are predicted by a Fully Convolutional Network\n(FCN), along with the joint confidence maps. To train our network, we collect a\nnew 3D human motion dataset capturing diverse total body motion of 40 subjects\nin a multiview system. We leverage a 3D deformable human model to reconstruct\ntotal body pose from the CNN outputs by exploiting the pose and shape prior in\nthe model. We also present a texture-based tracking method to obtain temporally\ncoherent motion capture output. We perform thorough quantitative evaluations\nincluding comparison with the existing body-specific and hand-specific methods,\nand performance analysis on camera viewpoint and human pose changes. Finally,\nwe demonstrate the results of our total body motion capture on various\nchallenging in-the-wild videos. Our code and newly collected human motion\ndataset will be publicly shared.","url_abs":"http://arxiv.org/abs/1812.01598v1","url_pdf":"http://arxiv.org/pdf/1812.01598v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"monocular-total-capture-posing-face-body-and","repo_url":"https://github.com/CMU-Perceptual-Computing-Lab/MonocularTotalCapture","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"hand-pose-estimation","task_name":"Hand Pose Estimation"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-3d-human-pose-estimation-on-human3","task":"Monocular 3D Human Pose Estimation","dataset":"Human3.6M","model":"Monocular Total Capture","rank_in_archive_order":30,"of":52,"metrics":{"Average MPJPE (mm)":"58.3","Frames Needed":"1","Need Ground Truth 2D Pose":"No","Use Video Sequence":"NO"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1812.01598","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}