{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/translate-to-recognize-networks-for-rgb-d","title":"Translate-to-Recognize Networks for RGB-D Scene Recognition","arxiv_id":"1904.12254","date":"2019-04-28","proceeding":"CVPR 2019 6","authors":["Dapeng Du","Li-Min Wang","Huiling Wang","Kai Zhao","Gangshan Wu"],"abstract":"Cross-modal transfer is helpful to enhance modality-specific discriminative\npower for scene recognition. To this end, this paper presents a unified\nframework to integrate the tasks of cross-modal translation and\nmodality-specific recognition, termed as Translate-to-Recognize Network\n(TRecgNet). Specifically, both translation and recognition tasks share the same\nencoder network, which allows to explicitly regularize the training of\nrecognition task with the help of translation, and thus improve its final\ngeneralization ability. For translation task, we place a decoder module on top\nof the encoder network and it is optimized with a new layer-wise semantic loss,\nwhile for recognition task, we use a linear classifier based on the feature\nembedding from encoder and its training is guided by the standard cross-entropy\nloss. In addition, our TRecgNet allows to exploit large numbers of unlabeled\nRGB-D data to train the translation task and thus improve the representation\npower of encoder network. Empirically, we verify that this new semi-supervised\nsetting is able to further enhance the performance of recognition network. We\nperform experiments on two RGB-D scene recognition benchmarks: NYU Depth v2 and\nSUN RGB-D, demonstrating that TRecgNet achieves superior performance to the\nexisting state-of-the-art methods, especially for recognition solely based on a\nsingle modality.","url_abs":"http://arxiv.org/abs/1904.12254v1","url_pdf":"http://arxiv.org/pdf/1904.12254v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"translate-to-recognize-networks-for-rgb-d","repo_url":"https://github.com/ownstyledu/Translate-to-Recognize-Networks","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.12254","atlas_url":"https://app.syntology.ai/?focus=1904.12254","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}