{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalizing-monocular-3d-human-pose","title":"Generalizing Monocular 3D Human Pose Estimation in the Wild","arxiv_id":"1904.05512","date":"2019-04-11","proceeding":null,"authors":["Luyang Wang","Yan Chen","Zhenhua Guo","Keyuan Qian","Mude Lin","Hongsheng Li","Jimmy S. Ren"],"abstract":"The availability of the large-scale labeled 3D poses in the Human3.6M dataset\nplays an important role in advancing the algorithms for 3D human pose\nestimation from a still image. We observe that recent innovation in this area\nmainly focuses on new techniques that explicitly address the generalization\nissue when using this dataset, because this database is constructed in a highly\ncontrolled environment with limited human subjects and background variations.\nDespite such efforts, we can show that the results of the current methods are\nstill error-prone especially when tested against the images taken in-the-wild.\nIn this paper, we aim to tackle this problem from a different perspective. We\npropose a principled approach to generate high quality 3D pose ground truth\ngiven any in-the-wild image with a person inside. We achieve this by first\ndevising a novel stereo inspired neural network to directly map any 2D pose to\nhigh quality 3D counterpart. We then perform a carefully designed geometric\nsearching scheme to further refine the joints. Based on this scheme, we build a\nlarge-scale dataset with 400,000 in-the-wild images and their corresponding 3D\npose ground truth. This enables the training of a high quality neural network\nmodel, without specialized training scheme and auxiliary loss function, which\nperforms favorably against the state-of-the-art 3D pose estimation methods. We\nalso evaluate the generalization ability of our model both quantitatively and\nqualitatively. Results show that our approach convincingly outperforms the\nprevious methods. We make our dataset and code publicly available.","url_abs":"http://arxiv.org/abs/1904.05512v1","url_pdf":"http://arxiv.org/pdf/1904.05512v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalizing-monocular-3d-human-pose","repo_url":"https://github.com/llcshappy/Monocular-3D-Human-Pose","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-human36m","task":"3D Human Pose Estimation","dataset":"Human3.6M","model":"Stereoscopic View Synthesis Subnetwork","rank_in_archive_order":80,"of":88,"metrics":{"Average MPJPE (mm)":"58","Multi-View or Monocular":"Multi-View","Using 2D ground-truth joints":"No"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-mpi-inf-3dhp","task":"3D Human Pose Estimation","dataset":"MPI-INF-3DHP","model":"Stereoscopic View Synthesis Subnetwork","rank_in_archive_order":101,"of":108,"metrics":{"AUC":"33.8","PCK":"71.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.05512","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}