{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/depth-cnns-for-rgb-d-scene-recognition","title":"Depth CNNs for RGB-D scene recognition: learning from scratch better than transferring from RGB-CNNs","arxiv_id":"1801.06797","date":"2018-01-21","proceeding":null,"authors":["Xinhang Song","Luis Herranz","Shuqiang Jiang"],"abstract":"Scene recognition with RGB images has been extensively studied and has\nreached very remarkable recognition levels, thanks to convolutional neural\nnetworks (CNN) and large scene datasets. In contrast, current RGB-D scene data\nis much more limited, so often leverages RGB large datasets, by transferring\npretrained RGB CNN models and fine-tuning with the target RGB-D dataset.\nHowever, we show that this approach has the limitation of hardly reaching\nbottom layers, which is key to learn modality-specific features. In contrast,\nwe focus on the bottom layers, and propose an alternative strategy to learn\ndepth features combining local weakly supervised training from patches followed\nby global fine tuning with images. This strategy is capable of learning very\ndiscriminative depth-specific features with limited depth images, without\nresorting to Places-CNN. In addition we propose a modified CNN architecture to\nfurther match the complexity of the model and the amount of data available. For\nRGB-D scene recognition, depth and RGB features are combined by projecting them\nin a common space and further leaning a multilayer classifier, which is jointly\noptimized in an end-to-end network. Our framework achieves state-of-the-art\naccuracy on NYU2 and SUN RGB-D in both depth only and combined RGB-D data.","url_abs":"http://arxiv.org/abs/1801.06797v1","url_pdf":"http://arxiv.org/pdf/1801.06797v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"depth-cnns-for-rgb-d-scene-recognition","repo_url":"https://github.com/songxinhang/D-CNN","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"scene-recognition","task_name":"Scene Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.06797","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}