{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/std2p-rgbd-semantic-segmentation-using-spatio","title":"STD2P: RGBD Semantic Segmentation Using Spatio-Temporal Data-Driven Pooling","arxiv_id":"1604.02388","date":"2016-04-08","proceeding":"CVPR 2017 7","authors":["Yang He","Wei-Chen Chiu","Margret Keuper","Mario Fritz"],"abstract":"We propose a novel superpixel-based multi-view convolutional neural network\nfor semantic image segmentation. The proposed network produces a high quality\nsegmentation of a single image by leveraging information from additional views\nof the same scene. Particularly in indoor videos such as captured by robotic\nplatforms or handheld and bodyworn RGBD cameras, nearby video frames provide\ndiverse viewpoints and additional context of objects and scenes. To leverage\nsuch information, we first compute region correspondences by optical flow and\nimage boundary-based superpixels. Given these region correspondences, we\npropose a novel spatio-temporal pooling layer to aggregate information over\nspace and time. We evaluate our approach on the NYU--Depth--V2 and the SUN3D\ndatasets and compare it to various state-of-the-art single-view and multi-view\napproaches. Besides a general improvement over the state-of-the-art, we also\nshow the benefits of making use of unlabeled frames during training for\nmulti-view as well as single-view prediction.","url_abs":"http://arxiv.org/abs/1604.02388v3","url_pdf":"http://arxiv.org/pdf/1604.02388v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"std2p-rgbd-semantic-segmentation-using-spatio","repo_url":"https://github.com/SSAW14/STD2P","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":null,"task_name":"RGBD Semantic Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"superpixels","task_name":"Superpixels"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-nyu-depth-v2","task":"Semantic Segmentation","dataset":"NYU Depth v2","model":"STD2P","rank_in_archive_order":110,"of":121,"metrics":{"Mean IoU":"40.1%"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}