{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hohonet-360-indoor-holistic-understanding","title":"HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features","arxiv_id":"2011.11498","date":"2020-11-23","proceeding":"CVPR 2021 1","authors":["Cheng Sun","Min Sun","Hwann-Tzong Chen"],"abstract":"We present HoHoNet, a versatile and efficient framework for holistic understanding of an indoor 360-degree panorama using a Latent Horizontal Feature (LHFeat). The compact LHFeat flattens the features along the vertical direction and has shown success in modeling per-column modality for room layout reconstruction. HoHoNet advances in two important aspects. First, the deep architecture is redesigned to run faster with improved accuracy. Second, we propose a novel horizon-to-dense module, which relaxes the per-column output shape constraint, allowing per-pixel dense prediction from LHFeat. HoHoNet is fast: It runs at 52 FPS and 110 FPS with ResNet-50 and ResNet-34 backbones respectively, for modeling dense modalities from a high-resolution $512 \\times 1024$ panorama. HoHoNet is also accurate. On the tasks of layout estimation and semantic segmentation, HoHoNet achieves results on par with current state-of-the-art. On dense depth estimation, HoHoNet outperforms all the prior arts by a large margin.","url_abs":"https://arxiv.org/abs/2011.11498v3","url_pdf":"https://arxiv.org/pdf/2011.11498v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hohonet-360-indoor-holistic-understanding","repo_url":"https://github.com/sunset1995/HoHoNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"3d-room-layouts-from-a-single-rgb-panorama","task_name":"3D Room Layouts From A Single RGB Panorama"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-room-layouts-from-a-single-rgb-panorama-on-3","task":"3D Room Layouts From A Single RGB Panorama","dataset":"Stanford2D3D Panoramic","model":"HoHoNet (ResNet-101)","rank_in_archive_order":6,"of":9,"metrics":{"3DIoU":"79.88"},"uses_additional_data":false},{"leaderboard":"/sota/depth-estimation-on-stanford2d3d-panoramic","task":"Depth Estimation","dataset":"Stanford2D3D Panoramic","model":"HoHoNet (ResNet-101)","rank_in_archive_order":14,"of":18,"metrics":{"RMSE":"0.3834","absolute relative error":"0.1014"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-stanford2d3d-1","task":"Semantic Segmentation","dataset":"Stanford2D3D Panoramic","model":"HoHoNet (ResNet-101)","rank_in_archive_order":13,"of":25,"metrics":{"mAcc":"65.0","mIoU":"52.0%"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-stanford2d3d-2","task":"Semantic Segmentation","dataset":"Stanford2D3D Panoramic - RGBD","model":"HoHoNet (ResNet-101)","rank_in_archive_order":3,"of":3,"metrics":{"mAcc":"68.9","mIoU":"56.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2011.11498","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}