{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/centerhmr-a-bottom-up-single-shot-method-for","title":"Monocular, One-stage, Regression of Multiple 3D People","arxiv_id":"2008.12272","date":"2020-08-27","proceeding":"ICCV 2021 10","authors":["Yu Sun","Qian Bao","Wu Liu","Yili Fu","Michael J. Black","Tao Mei"],"abstract":"This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a One-stage fashion for Multiple 3D People (termed ROMP). The approach is conceptually simple, bounding box-free, and able to learn a per-pixel representation in an end-to-end manner. Our method simultaneously predicts a Body Center heatmap and a Mesh Parameter map, which can jointly describe the 3D body mesh on the pixel level. Through a body-center-guided sampling process, the body mesh parameters of all people in the image are easily extracted from the Mesh Parameter map. Equipped with such a fine-grained representation, our one-stage framework is free of the complex multi-stage process and more robust to occlusion. Compared with state-of-the-art methods, ROMP achieves superior performance on the challenging multi-person benchmarks, including 3DPW and CMU Panoptic. Experiments on crowded/occluded datasets demonstrate the robustness under various types of occlusion. The released code is the first real-time implementation of monocular multi-person 3D mesh regression.","url_abs":"https://arxiv.org/abs/2008.12272v4","url_pdf":"https://arxiv.org/pdf/2008.12272v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"centerhmr-a-bottom-up-single-shot-method-for","repo_url":"https://github.com/Arthur151/ROMP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"centerhmr-a-bottom-up-single-shot-method-for","repo_url":"https://github.com/cai-jianfeng/ROMP_mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"mindspore","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"3d-depth-estimation","task_name":"3D Depth Estimation"},{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-multi-person-mesh-recovery","task_name":"3D Multi-Person Mesh Recovery"},{"task_slug":"3d-multi-person-pose-estimation","task_name":"3D Multi-Person Pose Estimation"},{"task_slug":"multi-person-pose-estimation","task_name":"Multi-Person Pose Estimation"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[{"method_slug":"heatmap","method_name":"Heatmap"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-depth-estimation-on-relative-human","task":"3D Depth Estimation","dataset":"Relative Human","model":"ROMP","rank_in_archive_order":2,"of":3,"metrics":{"PCDR":"54.84","PCDR-Adult":"55.34","PCDR-Baby":"30.08","PCDR-Kid":"48.41","PCDR-Teen":"51.12","mPCDK":"0.866"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-pose-estimation-on-3d-poses-in-the","task":"3D Human Pose Estimation","dataset":"3D Poses in the Wild Challenge","model":"ROMP","rank_in_archive_order":5,"of":5,"metrics":{"MPJPE":"81.76"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-emdb","task":"3D Human Pose Estimation","dataset":"EMDB","model":"ROMP","rank_in_archive_order":11,"of":13,"metrics":{"Average MPJAE (deg)":"26.5975","Average MPJAE-PA (deg)":"23.9901","Average MPJPE (mm)":"112.652","Average MPJPE-PA (mm)":"75.1869","Average MVE (mm)":"134.863","Average MVE-PA (mm)":"90.648","Jitter (10m/s^3)":"71.2556"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-cmu-panoptic","task":"3D Human Pose Estimation","dataset":"Panoptic","model":"ROMP (ResNet-50)","rank_in_archive_order":8,"of":9,"metrics":{"Average MPJPE (mm)":"127.6"},"uses_additional_data":false},{"leaderboard":"/sota/3d-multi-person-mesh-recovery-on-relative","task":"3D Multi-Person Mesh Recovery","dataset":"Relative Human","model":"ROMP","rank_in_archive_order":1,"of":1,"metrics":{"PCDR":"68.27"},"uses_additional_data":true},{"leaderboard":"/sota/multi-person-pose-estimation-on-crowdpose","task":"Multi-Person Pose Estimation","dataset":"CrowdPose","model":"ROMP+CAR","rank_in_archive_order":23,"of":28,"metrics":{"mAP @0.5:0.95":"58.6"},"uses_additional_data":false},{"leaderboard":"/sota/multi-person-pose-estimation-on-crowdpose","task":"Multi-Person Pose Estimation","dataset":"CrowdPose","model":"ROMP","rank_in_archive_order":27,"of":28,"metrics":{"mAP @0.5:0.95":"55.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2008.12272","atlas_url":"https://app.syntology.ai/?focus=2008.12272","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}