{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/back-to-optimization-diffusion-based-zero","title":"Back to Optimization: Diffusion-based Zero-Shot 3D Human Pose Estimation","arxiv_id":"2307.03833","date":"2023-07-07","proceeding":null,"authors":["Zhongyu Jiang","Zhuoran Zhou","Lei LI","Wenhao Chai","Cheng-Yen Yang","Jenq-Neng Hwang"],"abstract":"Learning-based methods have dominated the 3D human pose estimation (HPE) tasks with significantly better performance in most benchmarks than traditional optimization-based methods. Nonetheless, 3D HPE in the wild is still the biggest challenge for learning-based models, whether with 2D-3D lifting, image-to-3D, or diffusion-based methods, since the trained networks implicitly learn camera intrinsic parameters and domain-based 3D human pose distributions and estimate poses by statistical average. On the other hand, the optimization-based methods estimate results case-by-case, which can predict more diverse and sophisticated human poses in the wild. By combining the advantages of optimization-based and learning-based methods, we propose the \\textbf{Ze}ro-shot \\textbf{D}iffusion-based \\textbf{O}ptimization (\\textbf{ZeDO}) pipeline for 3D HPE to solve the problem of cross-domain and in-the-wild 3D HPE. Our multi-hypothesis \\textit{\\textbf{ZeDO}} achieves state-of-the-art (SOTA) performance on Human3.6M, with minMPJPE $51.4$mm, without training with any 2D-3D or image-3D pairs. Moreover, our single-hypothesis \\textit{\\textbf{ZeDO}} achieves SOTA performance on 3DPW dataset with PA-MPJPE $40.3$mm on cross-dataset evaluation, which even outperforms learning-based methods trained on 3DPW.","url_abs":"https://arxiv.org/abs/2307.03833v3","url_pdf":"https://arxiv.org/pdf/2307.03833v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"back-to-optimization-diffusion-based-zero","repo_url":"https://github.com/ipl-uw/ZeDO-Release","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"image-to-3d","task_name":"Image to 3D"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-3dpw","task":"3D Human Pose Estimation","dataset":"3DPW","model":"ZeDO (S=1,J=17)","rank_in_archive_order":65,"of":119,"metrics":{"MPJPE":"69.7","PA-MPJPE":"40.3"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-3dpw","task":"3D Human Pose Estimation","dataset":"3DPW","model":"ZeDO (Cross Dataset)","rank_in_archive_order":67,"of":119,"metrics":{"MPJPE":"80.9","PA-MPJPE":"42.6"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-mpi-inf-3dhp","task":"3D Human Pose Estimation","dataset":"MPI-INF-3DHP","model":"ZeDO (S=50)","rank_in_archive_order":23,"of":108,"metrics":{"AUC":"65.6","MPJPE":"55.2","PCK":"93"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.03833","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}