{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-multi-task-rgb-d-scene-analysis-for","title":"Efficient Multi-Task RGB-D Scene Analysis for Indoor Environments","arxiv_id":"2207.04526","date":"2022-07-10","proceeding":null,"authors":["Daniel Seichter","Söhnke Benedikt Fischedick","Mona Köhler","Horst-Michael Groß"],"abstract":"Semantic scene understanding is essential for mobile agents acting in various environments. Although semantic segmentation already provides a lot of information, details about individual objects as well as the general scene are missing but required for many real-world applications. However, solving multiple tasks separately is expensive and cannot be accomplished in real time given limited computing and battery capabilities on a mobile platform. In this paper, we propose an efficient multi-task approach for RGB-D scene analysis~(EMSANet) that simultaneously performs semantic and instance segmentation~(panoptic segmentation), instance orientation estimation, and scene classification. We show that all tasks can be accomplished using a single neural network in real time on a mobile platform without diminishing performance - by contrast, the individual tasks are able to benefit from each other. In order to evaluate our multi-task approach, we extend the annotations of the common RGB-D indoor datasets NYUv2 and SUNRGB-D for instance segmentation and orientation estimation. To the best of our knowledge, we are the first to provide results in such a comprehensive multi-task setting for indoor scene analysis on NYUv2 and SUNRGB-D.","url_abs":"https://arxiv.org/abs/2207.04526v1","url_pdf":"https://arxiv.org/pdf/2207.04526v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-multi-task-rgb-d-scene-analysis-for","repo_url":"https://github.com/tui-nicr/emsanet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"efficient-multi-task-rgb-d-scene-analysis-for","repo_url":"https://github.com/tui-nicr/nicr-scene-analysis-datasets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"panoptic-segmentation","task_name":"Panoptic Segmentation"},{"task_slug":"scene-classification","task_name":"Scene Classification"},{"task_slug":"scene-classification-unified-classes","task_name":"Scene Classification (unified classes)"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/panoptic-segmentation-on-nyu-depth-v2","task":"Panoptic Segmentation","dataset":"NYU Depth v2","model":"EMSANet","rank_in_archive_order":2,"of":2,"metrics":{"PQ":"47.38"},"uses_additional_data":false},{"leaderboard":"/sota/panoptic-segmentation-on-sun-rgbd","task":"Panoptic Segmentation","dataset":"SUN-RGBD","model":"EMSANet","rank_in_archive_order":1,"of":1,"metrics":{"PQ":"52.84"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-nyu-depth-v2","task":"Semantic Segmentation","dataset":"NYU Depth v2","model":"EMSANet (2x ResNet-34 NBt1D, finetuned)","rank_in_archive_order":37,"of":121,"metrics":{"Mean IoU":"53.34%"},"uses_additional_data":true},{"leaderboard":"/sota/semantic-segmentation-on-sun-rgbd","task":"Semantic Segmentation","dataset":"SUN-RGBD","model":"DPLNet","rank_in_archive_order":30,"of":44,"metrics":{"Mean IoU":"48.47%"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2207.04526","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}