{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-3d-pose-estimation-with","title":"Unsupervised 3D Pose Estimation with Geometric Self-Supervision","arxiv_id":"1904.04812","date":"2019-04-09","proceeding":"CVPR 2019 6","authors":["Ching-Hang Chen","Ambrish Tyagi","Amit Agrawal","Dylan Drover","Rohith MV","Stefan Stojanov","James M. Rehg"],"abstract":"We present an unsupervised learning approach to recover 3D human pose from 2D\nskeletal joints extracted from a single image. Our method does not require any\nmulti-view image data, 3D skeletons, correspondences between 2D-3D points, or\nuse previously learned 3D priors during training. A lifting network accepts 2D\nlandmarks as inputs and generates a corresponding 3D skeleton estimate. During\ntraining, the recovered 3D skeleton is reprojected on random camera viewpoints\nto generate new \"synthetic\" 2D poses. By lifting the synthetic 2D poses back to\n3D and re-projecting them in the original camera view, we can define\nself-consistency loss both in 3D and in 2D. The training can thus be self\nsupervised by exploiting the geometric self-consistency of the\nlift-reproject-lift process. We show that self-consistency alone is not\nsufficient to generate realistic skeletons, however adding a 2D pose\ndiscriminator enables the lifter to output valid 3D poses. Additionally, to\nlearn from 2D poses \"in the wild\", we train an unsupervised 2D domain adapter\nnetwork to allow for an expansion of 2D data. This improves results and\ndemonstrates the usefulness of 2D pose data for unsupervised 3D lifting.\nResults on Human3.6M dataset for 3D human pose estimation demonstrate that our\napproach improves upon the previous unsupervised methods by 30% and outperforms\nmany weakly supervised approaches that explicitly use 3D data.","url_abs":"http://arxiv.org/abs/1904.04812v1","url_pdf":"http://arxiv.org/pdf/1904.04812v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":null,"task_name":"valid"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-mpi-inf-3dhp","task":"3D Human Pose Estimation","dataset":"MPI-INF-3DHP","model":"2D-3D Lifting Network","rank_in_archive_order":99,"of":108,"metrics":{"AUC":"36.3","PCK":"71.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.04812","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}