{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sparseness-meets-deepness-3d-human-pose","title":"Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video","arxiv_id":"1511.09439","date":"2015-11-30","proceeding":"CVPR 2016 6","authors":["Xiaowei Zhou","Menglong Zhu","Spyridon Leonardos","Kosta Derpanis","Kostas Daniilidis"],"abstract":"This paper addresses the challenge of 3D full-body human pose estimation from\na monocular image sequence. Here, two cases are considered: (i) the image\nlocations of the human joints are provided and (ii) the image locations of\njoints are unknown. In the former case, a novel approach is introduced that\nintegrates a sparsity-driven 3D geometric prior and temporal smoothness. In the\nlatter case, the former case is extended by treating the image locations of the\njoints as latent variables. A deep fully convolutional network is trained to\npredict the uncertainty maps of the 2D joint locations. The 3D pose estimates\nare realized via an Expectation-Maximization algorithm over the entire\nsequence, where it is shown that the 2D joint location uncertainties can be\nconveniently marginalized out during inference. Empirical evaluation on the\nHuman3.6M dataset shows that the proposed approaches achieve greater 3D pose\nestimation accuracy over state-of-the-art baselines. Further, the proposed\napproach outperforms a publicly available 2D pose estimation baseline on the\nchallenging PennAction dataset.","url_abs":"http://arxiv.org/abs/1511.09439v2","url_pdf":"http://arxiv.org/pdf/1511.09439v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sparseness-meets-deepness-3d-human-pose","repo_url":"https://github.com/chuxiaoselena/SparsenessMeetsDeepness","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"torch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"2d-pose-estimation","task_name":"2D Pose Estimation"},{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-3d-human-pose-estimation-on-human3","task":"Monocular 3D Human Pose Estimation","dataset":"Human3.6M","model":"Sparseness Meets Deepness","rank_in_archive_order":40,"of":52,"metrics":{"Average MPJPE (mm)":"113.01","Frames Needed":"300","Need Ground Truth 2D Pose":"No","Use Video Sequence":"Yes"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1511.09439","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}