{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-generic-diffusion-based-approach-for-3d","title":"A generic diffusion-based approach for 3D human pose prediction in the wild","arxiv_id":"2210.05669","date":"2022-10-11","proceeding":null,"authors":["Saeed Saadatnejad","Ali Rasekh","Mohammadreza Mofayezi","Yasamin Medghalchi","Sara Rajabzadeh","Taylor Mordan","Alexandre Alahi"],"abstract":"Predicting 3D human poses in real-world scenarios, also known as human pose forecasting, is inevitably subject to noisy inputs arising from inaccurate 3D pose estimations and occlusions. To address these challenges, we propose a diffusion-based approach that can predict given noisy observations. We frame the prediction task as a denoising problem, where both observation and prediction are considered as a single sequence containing missing elements (whether in the observation or prediction horizon). All missing elements are treated as noise and denoised with our conditional diffusion model. To better handle long-term forecasting horizon, we present a temporal cascaded diffusion model. We demonstrate the benefits of our approach on four publicly available datasets (Human3.6M, HumanEva-I, AMASS, and 3DPW), outperforming the state-of-the-art. Additionally, we show that our framework is generic enough to improve any 3D pose prediction model as a pre-processing step to repair their inputs and a post-processing step to refine their outputs. The code is available online: \\url{https://github.com/vita-epfl/DePOSit}.","url_abs":"https://arxiv.org/abs/2210.05669v2","url_pdf":"https://arxiv.org/pdf/2210.05669v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-generic-diffusion-based-approach-for-3d","repo_url":"https://github.com/vita-epfl/deposit","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"AGPL-3.0"}}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"human-pose-forecasting","task_name":"Human Pose Forecasting"},{"task_slug":"missing-elements","task_name":"Missing Elements"},{"task_slug":"pose-prediction","task_name":"Pose Prediction"},{"task_slug":"prediction","task_name":"Prediction"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"},{"method_slug":"repair","method_name":"Repair"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-pose-forecasting-on-3dpw","task":"Human Pose Forecasting","dataset":"3DPW","model":"TCD","rank_in_archive_order":7,"of":7,"metrics":{"FDE@1000ms (mm)":"73.4","FDE@560ms (mm)":"55.4","FDE@720ms (mm)":"61.6","FDE@880ms (mm)":"67.9"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-amass","task":"Human Pose Forecasting","dataset":"AMASS","model":"TCD","rank_in_archive_order":6,"of":11,"metrics":{"FDE@1000ms (mm)":"66.7","FDE@560ms (mm)":"49.8","FDE@720ms (mm)":"54.5","FDE@880ms (mm)":"60.1"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-human36m","task":"Human Pose Forecasting","dataset":"Human3.6M","model":"TCD","rank_in_archive_order":22,"of":33,"metrics":{"ADE":"356","APD":"19466","FDE":"396","MMADE":"463","MMFDE":"445"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-humaneva-i","task":"Human Pose Forecasting","dataset":"HumanEva-I","model":"TCD","rank_in_archive_order":1,"of":11,"metrics":{"ADE@2000ms":"199","APD@2000ms":"6764","FDE@2000ms":"215"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.05669","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}