{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-supervised-learning-of-motion-capture","title":"Self-supervised Learning of Motion Capture","arxiv_id":"1712.01337","date":"2017-12-04","proceeding":"NeurIPS 2017 12","authors":["Hsiao-Yu Fish Tung","Hsiao-Wei Tung","Ersin Yumer","Katerina Fragkiadaki"],"abstract":"Current state-of-the-art solutions for motion capture from a single camera\nare optimization driven: they optimize the parameters of a 3D human model so\nthat its re-projection matches measurements in the video (e.g. person\nsegmentation, optical flow, keypoint detections etc.). Optimization models are\nsusceptible to local minima. This has been the bottleneck that forced using\nclean green-screen like backgrounds at capture time, manual initialization, or\nswitching to multiple cameras as input resource. In this work, we propose a\nlearning based motion capture model for single camera input. Instead of\noptimizing mesh and skeleton parameters directly, our model optimizes neural\nnetwork weights that predict 3D shape and skeleton configurations given a\nmonocular RGB video. Our model is trained using a combination of strong\nsupervision from synthetic data, and self-supervision from differentiable\nrendering of (a) skeletal keypoints, (b) dense 3D mesh motion, and (c)\nhuman-background segmentation, in an end-to-end framework. Empirically we show\nour model combines the best of both worlds of supervised learning and test-time\noptimization: supervised learning initializes the model parameters in the right\nregime, ensuring good pose and surface initialization at test time, without\nmanual effort. Self-supervision by back-propagating through differentiable\nrendering allows (unsupervised) adaptation of the model to the test data, and\noffers much tighter fit than a pretrained fixed model. We show that the\nproposed model improves with experience and converges to low-error solutions\nwhere previous optimization methods fail.","url_abs":"http://arxiv.org/abs/1712.01337v1","url_pdf":"http://arxiv.org/pdf/1712.01337v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-supervised-learning-of-motion-capture","repo_url":"https://github.com/chingswy/HumanPoseMemo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-human-reconstruction","task_name":"3D Human Reconstruction"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"weakly-supervised-3d-human-pose-estimation","task_name":"Weakly-supervised 3D Human Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-surreal-1","task":"3D Human Pose Estimation","dataset":"Surreal","model":"self-supervised mocap","rank_in_archive_order":5,"of":6,"metrics":{"MPJPE":"64.4"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-reconstruction-on-surreal","task":"3D Human Reconstruction","dataset":"Surreal","model":"self-supervised mocap","rank_in_archive_order":2,"of":2,"metrics":{"MPVPE":"74.5"},"uses_additional_data":false},{"leaderboard":"/sota/weakly-supervised-3d-human-pose-estimation-on","task":"Weakly-supervised 3D Human Pose Estimation","dataset":"Human3.6M","model":"self-supervised mocap","rank_in_archive_order":25,"of":33,"metrics":{"Average MPJPE (mm)":"98.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.01337","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}