{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/genpose-generative-category-level-object-pose","title":"GenPose: Generative Category-level Object Pose Estimation via Diffusion Models","arxiv_id":"2306.10531","date":"2023-06-18","proceeding":null,"authors":["Jiyao Zhang","Mingdong Wu","Hao Dong"],"abstract":"Object pose estimation plays a vital role in embodied AI and computer vision, enabling intelligent agents to comprehend and interact with their surroundings. Despite the practicality of category-level pose estimation, current approaches encounter challenges with partially observed point clouds, known as the multihypothesis issue. In this study, we propose a novel solution by reframing categorylevel object pose estimation as conditional generative modeling, departing from traditional point-to-point regression. Leveraging score-based diffusion models, we estimate object poses by sampling candidates from the diffusion model and aggregating them through a two-step process: filtering out outliers via likelihood estimation and subsequently mean-pooling the remaining candidates. To avoid the costly integration process when estimating the likelihood, we introduce an alternative method that trains an energy-based model from the original score-based model, enabling end-to-end likelihood estimation. Our approach achieves state-of-the-art performance on the REAL275 dataset, surpassing 50% and 60% on strict 5d2cm and 5d5cm metrics, respectively. Furthermore, our method demonstrates strong generalizability to novel categories sharing similar symmetric properties without fine-tuning and can readily adapt to object pose tracking tasks, yielding comparable results to the current state-of-the-art baselines.","url_abs":"https://arxiv.org/abs/2306.10531v3","url_pdf":"https://arxiv.org/pdf/2306.10531v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"6d-pose-estimation-1","task_name":"6D Pose Estimation"},{"task_slug":"6d-pose-estimation-using-rgbd","task_name":"6D Pose Estimation using RGBD"},{"task_slug":"object","task_name":"Object"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"pose-tracking","task_name":"Pose Tracking"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/6d-pose-estimation-using-rgbd-on-real275","task":"6D Pose Estimation using RGBD","dataset":"REAL275","model":"GenPose https://github.com/Jiyao06/GenPose","rank_in_archive_order":1,"of":11,"metrics":{"mAP 10, 2cm":"72.4","mAP 10, 5cm":"84.0","mAP 5, 2cm":"52.1","mAP 5, 5cm":"60.9"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2306.10531","atlas_url":"https://app.syntology.ai/?focus=2306.10531","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}