{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/expressive-body-capture-3d-hands-face-and","title":"Expressive Body Capture: 3D Hands, Face, and Body from a Single Image","arxiv_id":"1904.05866","date":"2019-04-11","proceeding":"CVPR 2019 6","authors":["Georgios Pavlakos","Vasileios Choutas","Nima Ghorbani","Timo Bolkart","Ahmed A. A. Osman","Dimitrios Tzionas","Michael J. Black"],"abstract":"To facilitate the analysis of human actions, interactions and emotions, we\ncompute a 3D model of human body pose, hand pose, and facial expression from a\nsingle monocular image. To achieve this, we use thousands of 3D scans to train\na new, unified, 3D model of the human body, SMPL-X, that extends SMPL with\nfully articulated hands and an expressive face. Learning to regress the\nparameters of SMPL-X directly from images is challenging without paired images\nand 3D ground truth. Consequently, we follow the approach of SMPLify, which\nestimates 2D features and then optimizes model parameters to fit the features.\nWe improve on SMPLify in several significant ways: (1) we detect 2D features\ncorresponding to the face, hands, and feet and fit the full SMPL-X model to\nthese; (2) we train a new neural network pose prior using a large MoCap\ndataset; (3) we define a new interpenetration penalty that is both fast and\naccurate; (4) we automatically detect gender and the appropriate body models\n(male, female, or neutral); (5) our PyTorch implementation achieves a speedup\nof more than 8x over Chumpy. We use the new method, SMPLify-X, to fit SMPL-X to\nboth controlled images and images in the wild. We evaluate 3D accuracy on a new\ncurated dataset comprising 100 images with pseudo ground-truth. This is a step\ntowards automatic expressive human capture from monocular RGB data. The models,\ncode, and data are available for research purposes at\nhttps://smpl-x.is.tue.mpg.de.","url_abs":"http://arxiv.org/abs/1904.05866v1","url_pdf":"http://arxiv.org/pdf/1904.05866v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"expressive-body-capture-3d-hands-face-and","repo_url":"https://github.com/vchoutas/smplify-x","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-human-reconstruction","task_name":"3D Human Reconstruction"},{"task_slug":"3d-multi-person-mesh-recovery","task_name":"3D Multi-Person Mesh Recovery"},{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-reconstruction-on-agora-1","task":"3D Human Reconstruction","dataset":"AGORA","model":"SMPLify-X","rank_in_archive_order":4,"of":5,"metrics":{"B-MPJPE":"182.1","B-MVE":"187.0","B-NMJE":"256.5","B-NMVE":"263.3","F-MPJPE":"52.9","F-MVE":"48.9","FB-MPJPE":"231.8","FB-MVE":"236.5","FB-NMJE":"326.5","FB-NMVE":"333.1","LH/RH-MPJPE":"46.5/49.6","LH/RH-MVE":"48.3/51.4"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-reconstruction-on-expressive-hands-1","task":"3D Human Reconstruction","dataset":"Expressive hands and faces dataset (EHF)","model":"SMPLify-X","rank_in_archive_order":5,"of":5,"metrics":{"MPJPE, left hand":"12.2","MPJPE-14":"87.6","PA V2V (mm), body only":"75.4","PA V2V (mm), face":"4.9","PA V2V (mm), left hand":"11.6","TR V2V (mm), body only":"116.1","TR V2V (mm), face":"11.5","TR V2V (mm), left hand":"23.8","TR V2V (mm), whole body":"93.0","mean P2S":"36.8","median P2S":"23.0"},"uses_additional_data":false},{"leaderboard":"/sota/3d-multi-person-mesh-recovery-on-agora","task":"3D Multi-Person Mesh Recovery","dataset":"AGORA","model":"SMPLify-X","rank_in_archive_order":7,"of":7,"metrics":{"B-MPJPE":"182.1","B-MVE":"187.0","B-NMJE":"256.5","B-NMVE":"263.3","F-MPJPE":"52.9","F-MVE":"48.9","FB-MPJPE":"231.8","FB-MVE":"236.5","FB-NMJE":"326.5","FB-NMVE":"333.1","LH/RH-MPJPE":"46.5/49.6","LH/RH-MVE":"48.3/51.4"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.05866","atlas_url":"https://app.syntology.ai/?focus=1904.05866","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}