{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/compressed-volumetric-heatmaps-for-multi","title":"Compressed Volumetric Heatmaps for Multi-Person 3D Pose Estimation","arxiv_id":"2004.00329","date":"2020-04-01","proceeding":"CVPR 2020 6","authors":["Matteo Fabbri","Fabio Lanzi","Simone Calderara","Stefano Alletto","Rita Cucchiara"],"abstract":"In this paper we present a novel approach for bottom-up multi-person 3D human pose estimation from monocular RGB images. We propose to use high resolution volumetric heatmaps to model joint locations, devising a simple and effective compression method to drastically reduce the size of this representation. At the core of the proposed method lies our Volumetric Heatmap Autoencoder, a fully-convolutional network tasked with the compression of ground-truth heatmaps into a dense intermediate representation. A second model, the Code Predictor, is then trained to predict these codes, which can be decompressed at test time to re-obtain the original representation. Our experimental evaluation shows that our method performs favorably when compared to state of the art on both multi-person and single-person 3D human pose estimation datasets and, thanks to our novel compression strategy, can process full-HD images at the constant runtime of 8 fps regardless of the number of subjects in the scene. Code and models available at https://github.com/fabbrimatteo/LoCO .","url_abs":"https://arxiv.org/abs/2004.00329v1","url_pdf":"https://arxiv.org/pdf/2004.00329v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"compressed-volumetric-heatmaps-for-multi","repo_url":"https://github.com/fabbrimatteo/LoCO","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"heatmap","method_name":"Heatmap"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-human36m","task":"3D Human Pose Estimation","dataset":"Human3.6M","model":"LoCO","rank_in_archive_order":71,"of":88,"metrics":{"Average MPJPE (mm)":"51.1","Multi-View or Monocular":"Monocular","PA-MPJPE":"43.4","Using 2D ground-truth joints":"No"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-cmu-panoptic","task":"3D Human Pose Estimation","dataset":"Panoptic","model":"LoCO","rank_in_archive_order":6,"of":9,"metrics":{"Average MPJPE (mm)":"69"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2004.00329","atlas_url":"https://app.syntology.ai/?focus=2004.00329","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}