{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recurrent-scene-parsing-with-perspective","title":"Recurrent Scene Parsing with Perspective Understanding in the Loop","arxiv_id":"1705.07238","date":"2017-05-20","proceeding":"CVPR 2018 6","authors":["Shu Kong","Charless Fowlkes"],"abstract":"Objects may appear at arbitrary scales in perspective images of a scene,\nposing a challenge for recognition systems that process images at a fixed\nresolution. We propose a depth-aware gating module that adaptively selects the\npooling field size in a convolutional network architecture according to the\nobject scale (inversely proportional to the depth) so that small details are\npreserved for distant objects while larger receptive fields are used for those\nnearby. The depth gating signal is provided by stereo disparity or estimated\ndirectly from monocular input. We integrate this depth-aware gating into a\nrecurrent convolutional neural network to perform semantic segmentation. Our\nrecurrent module iteratively refines the segmentation results, leveraging the\ndepth and semantic predictions from the previous iterations.\n  Through extensive experiments on four popular large-scale RGB-D datasets, we\ndemonstrate this approach achieves competitive semantic segmentation\nperformance with a model which is substantially more compact. We carry out\nextensive analysis of this architecture including variants that operate on\nmonocular RGB but use depth as side-information during training, unsupervised\ngating as a generic attentional mechanism, and multi-resolution gating. We find\nthat gated pooling for joint semantic segmentation and depth yields\nstate-of-the-art results for quantitative monocular depth estimation.","url_abs":"http://arxiv.org/abs/1705.07238v2","url_pdf":"http://arxiv.org/pdf/1705.07238v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recurrent-scene-parsing-with-perspective","repo_url":"https://github.com/aimerykong/Recurrent-Scene-Parsing-with-Perspective-Understanding-in-the-loop","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"scene-parsing","task_name":"Scene Parsing"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-cityscapes","task":"Semantic Segmentation","dataset":"Cityscapes test","model":"DepthSeg (ResNet-101)","rank_in_archive_order":61,"of":105,"metrics":{"Mean IoU (class)":"78.2%"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-nyu-depth-v2","task":"Semantic Segmentation","dataset":"NYU Depth v2","model":"RecurrentSceneParsing","rank_in_archive_order":96,"of":121,"metrics":{"Mean IoU":"44.5%"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-sun-rgbd","task":"Semantic Segmentation","dataset":"SUN-RGBD","model":"DPLNet","rank_in_archive_order":40,"of":44,"metrics":{"Mean IoU":"45.1%"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.07238","atlas_url":"https://app.syntology.ai/?focus=1705.07238","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}