{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monoperfcap-human-performance-capture-from","title":"MonoPerfCap: Human Performance Capture from Monocular Video","arxiv_id":"1708.02136","date":"2017-08-07","proceeding":null,"authors":["Weipeng Xu","Avishek Chatterjee","Michael Zollhöfer","Helge Rhodin","Dushyant Mehta","Hans-Peter Seidel","Christian Theobalt"],"abstract":"We present the first marker-less approach for temporally coherent 3D\nperformance capture of a human with general clothing from monocular video. Our\napproach reconstructs articulated human skeleton motion as well as medium-scale\nnon-rigid surface deformations in general scenes. Human performance capture is\na challenging problem due to the large range of articulation, potentially fast\nmotion, and considerable non-rigid deformations, even from multi-view data.\nReconstruction from monocular video alone is drastically more challenging,\nsince strong occlusions and the inherent depth ambiguity lead to a highly\nill-posed reconstruction problem. We tackle these challenges by a novel\napproach that employs sparse 2D and 3D human pose detections from a\nconvolutional neural network using a batch-based pose estimation strategy.\nJoint recovery of per-batch motion allows to resolve the ambiguities of the\nmonocular reconstruction problem based on a low dimensional trajectory\nsubspace. In addition, we propose refinement of the surface geometry based on\nfully automatically extracted silhouettes to enable medium-scale non-rigid\nalignment. We demonstrate state-of-the-art performance capture results that\nenable exciting applications such as video editing and free viewpoint video,\npreviously infeasible from monocular video. Our qualitative and quantitative\nevaluation demonstrates that our approach significantly outperforms previous\nmonocular methods in terms of accuracy, robustness and scene complexity that\ncan be handled.","url_abs":"http://arxiv.org/abs/1708.02136v2","url_pdf":"http://arxiv.org/pdf/1708.02136v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"monocular-reconstruction","task_name":"Monocular Reconstruction"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"video-editing","task_name":"Video Editing"}],"methods":[],"datasets_introduced":[{"slug":"monoperfcap-dataset","name":"MonoPerfCap Dataset","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}