{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/harvesting-multiple-views-for-marker-less-3d","title":"Harvesting Multiple Views for Marker-less 3D Human Pose Annotations","arxiv_id":"1704.04793","date":"2017-04-16","proceeding":"CVPR 2017 7","authors":["Georgios Pavlakos","Xiaowei Zhou","Konstantinos G. Derpanis","Kostas Daniilidis"],"abstract":"Recent advances with Convolutional Networks (ConvNets) have shifted the\nbottleneck for many computer vision tasks to annotated data collection. In this\npaper, we present a geometry-driven approach to automatically collect\nannotations for human pose prediction tasks. Starting from a generic ConvNet\nfor 2D human pose, and assuming a multi-view setup, we describe an automatic\nway to collect accurate 3D human pose annotations. We capitalize on constraints\noffered by the 3D geometry of the camera setup and the 3D structure of the\nhuman body to probabilistically combine per view 2D ConvNet predictions into a\nglobally optimal 3D pose. This 3D pose is used as the basis for harvesting\nannotations. The benefit of the annotations produced automatically with our\napproach is demonstrated in two challenging settings: (i) fine-tuning a generic\nConvNet-based 2D pose predictor to capture the discriminative aspects of a\nsubject's appearance (i.e.,\"personalization\"), and (ii) training a ConvNet from\nscratch for single view 3D human pose prediction without leveraging 3D pose\ngroundtruth. The proposed multi-view pose estimator achieves state-of-the-art\nresults on standard benchmarks, demonstrating the effectiveness of our method\nin exploiting the available multi-view information.","url_abs":"http://arxiv.org/abs/1704.04793v1","url_pdf":"http://arxiv.org/pdf/1704.04793v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-geometry","task_name":"3D geometry"},{"task_slug":"pose-prediction","task_name":"Pose Prediction"},{"task_slug":"weakly-supervised-3d-human-pose-estimation","task_name":"Weakly-supervised 3D Human Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/weakly-supervised-3d-human-pose-estimation-on","task":"Weakly-supervised 3D Human Pose Estimation","dataset":"Human3.6M","model":"Pavlakos et al.","rank_in_archive_order":28,"of":33,"metrics":{"3D Annotations":"No","Average MPJPE (mm)":"118.4"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1704.04793","atlas_url":"https://app.syntology.ai/?focus=1704.04793","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}