{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-3d-human-pose-estimation-in-the-wild","title":"Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach","arxiv_id":"1704.02447","date":"2017-04-08","proceeding":"ICCV 2017 10","authors":["Xingyi Zhou","Qi-Xing Huang","Xiao Sun","xiangyang xue","Yichen Wei"],"abstract":"In this paper, we study the task of 3D human pose estimation in the wild.\nThis task is challenging due to lack of training data, as existing datasets are\neither in the wild images with 2D pose or in the lab images with 3D pose.\n  We propose a weakly-supervised transfer learning method that uses mixed 2D\nand 3D labels in a unified deep neutral network that presents two-stage\ncascaded structure. Our network augments a state-of-the-art 2D pose estimation\nsub-network with a 3D depth regression sub-network. Unlike previous two stage\napproaches that train the two sub-networks sequentially and separately, our\ntraining is end-to-end and fully exploits the correlation between the 2D pose\nand depth estimation sub-tasks. The deep features are better learnt through\nshared representations. In doing so, the 3D pose labels in controlled lab\nenvironments are transferred to in the wild images. In addition, we introduce a\n3D geometric constraint to regularize the 3D pose prediction, which is\neffective in the absence of ground truth depth labels. Our method achieves\ncompetitive results on both 2D and 3D benchmarks.","url_abs":"http://arxiv.org/abs/1704.02447v2","url_pdf":"http://arxiv.org/pdf/1704.02447v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-3d-human-pose-estimation-in-the-wild","repo_url":"https://github.com/xingyizhou/pose-hg-3d","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"towards-3d-human-pose-estimation-in-the-wild","repo_url":"https://github.com/ECE740F21T01/pytorch-pose-hg-3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"towards-3d-human-pose-estimation-in-the-wild","repo_url":"https://github.com/mengyingfei/pose-3d-pytorch-ros","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"towards-3d-human-pose-estimation-in-the-wild","repo_url":"https://github.com/mikeshihyaolin/pose-hg-3d-preprocessing","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"towards-3d-human-pose-estimation-in-the-wild","repo_url":"https://github.com/nish-97v/3D-human-pose-estimation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"towards-3d-human-pose-estimation-in-the-wild","repo_url":"https://github.com/xingyizhou/pytorch-pose-hg-3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"2d-pose-estimation","task_name":"2D Pose Estimation"},{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-multi-person-pose-estimation-absolute","task_name":"3D Multi-Person Pose Estimation (absolute)"},{"task_slug":"3d-multi-person-pose-estimation-root-relative","task_name":"3D Multi-Person Pose Estimation (root-relative)"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"pose-prediction","task_name":"Pose Prediction"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-geometric-pose","task":"3D Human Pose Estimation","dataset":"Geometric Pose Affordance","model":"Baseline model","rank_in_archive_order":1,"of":1,"metrics":{"MPJPE (CA)":"89.2","MPJPE (CS)":"99.4","PCK3D (CA)":"83.6","PCK3D (CS)":"81.3"},"uses_additional_data":false},{"leaderboard":"/sota/monocular-3d-human-pose-estimation-on-human3","task":"Monocular 3D Human Pose Estimation","dataset":"Human3.6M","model":"Weakly Supervised Transfer Learning","rank_in_archive_order":34,"of":52,"metrics":{"Average MPJPE (mm)":"64.9","Frames Needed":"1","Need Ground Truth 2D Pose":"No","Use Video Sequence":"No"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.02447","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}