{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ordinal-depth-supervision-for-3d-human-pose","title":"Ordinal Depth Supervision for 3D Human Pose Estimation","arxiv_id":"1805.04095","date":"2018-05-10","proceeding":"CVPR 2018 6","authors":["Georgios Pavlakos","Xiaowei Zhou","Kostas Daniilidis"],"abstract":"Our ability to train end-to-end systems for 3D human pose estimation from\nsingle images is currently constrained by the limited availability of 3D\nannotations for natural images. Most datasets are captured using Motion Capture\n(MoCap) systems in a studio setting and it is difficult to reach the\nvariability of 2D human pose datasets, like MPII or LSP. To alleviate the need\nfor accurate 3D ground truth, we propose to use a weaker supervision signal\nprovided by the ordinal depths of human joints. This information can be\nacquired by human annotators for a wide range of images and poses. We showcase\nthe effectiveness and flexibility of training Convolutional Networks (ConvNets)\nwith these ordinal relations in different settings, always achieving\ncompetitive performance with ConvNets trained with accurate 3D joint\ncoordinates. Additionally, to demonstrate the potential of the approach, we\naugment the popular LSP and MPII datasets with ordinal depth annotations. This\nextension allows us to present quantitative and qualitative evaluation in\nnon-studio conditions. Simultaneously, these ordinal annotations can be easily\nincorporated in the training procedure of typical ConvNets for 3D human pose.\nThrough this inclusion we achieve new state-of-the-art performance for the\nrelevant benchmarks and validate the effectiveness of ordinal depth supervision\nfor 3D human pose.","url_abs":"http://arxiv.org/abs/1805.04095v1","url_pdf":"http://arxiv.org/pdf/1805.04095v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ordinal-depth-supervision-for-3d-human-pose","repo_url":"https://github.com/geopavlakos/ordinal-pose3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-humaneva-i","task":"3D Human Pose Estimation","dataset":"HumanEva-I","model":"Ordinal Depth Supervision","rank_in_archive_order":8,"of":31,"metrics":{"Mean Reconstruction Error (mm)":"18.3"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-mpi-inf-3dhp","task":"3D Human Pose Estimation","dataset":"MPI-INF-3DHP","model":"Ordinal Depth Supervision","rank_in_archive_order":100,"of":108,"metrics":{"AUC":"35.3","PCK":"71.9"},"uses_additional_data":false},{"leaderboard":"/sota/monocular-3d-human-pose-estimation-on-human3","task":"Monocular 3D Human Pose Estimation","dataset":"Human3.6M","model":"Ordinal Depth Supervision","rank_in_archive_order":44,"of":52,"metrics":{"Frames Needed":"1","Need Ground Truth 2D Pose":"No","Use Video Sequence":"No"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.04095","atlas_url":"https://app.syntology.ai/?focus=1805.04095","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}