{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/knowledge-guided-disambiguation-for-large","title":"Knowledge Guided Disambiguation for Large-Scale Scene Classification with Multi-Resolution CNNs","arxiv_id":"1610.01119","date":"2016-10-04","proceeding":null,"authors":["Limin Wang","Sheng Guo","Weilin Huang","Yuanjun Xiong","Yu Qiao"],"abstract":"Convolutional Neural Networks (CNNs) have made remarkable progress on scene\nrecognition, partially due to these recent large-scale scene datasets, such as\nthe Places and Places2. Scene categories are often defined by multi-level\ninformation, including local objects, global layout, and background\nenvironment, thus leading to large intra-class variations. In addition, with\nthe increasing number of scene categories, label ambiguity has become another\ncrucial issue in large-scale classification. This paper focuses on large-scale\nscene recognition and makes two major contributions to tackle these issues.\nFirst, we propose a multi-resolution CNN architecture that captures visual\ncontent and structure at multiple levels. The multi-resolution CNNs are\ncomposed of coarse resolution CNNs and fine resolution CNNs, which are\ncomplementary to each other. Second, we design two knowledge guided\ndisambiguation techniques to deal with the problem of label ambiguity. (i) We\nexploit the knowledge from the confusion matrix computed on validation data to\nmerge ambiguous classes into a super category. (ii) We utilize the knowledge of\nextra networks to produce a soft label for each image. Then the super\ncategories or soft labels are employed to guide CNN training on the Places2. We\nconduct extensive experiments on three large-scale image datasets (ImageNet,\nPlaces, and Places2), demonstrating the effectiveness of our approach.\nFurthermore, our method takes part in two major scene recognition challenges,\nand achieves the second place at the Places2 challenge in ILSVRC 2015, and the\nfirst place at the LSUN challenge in CVPR 2016. Finally, we directly test the\nlearned representations on other scene benchmarks, and obtain the new\nstate-of-the-art results on the MIT Indoor67 (86.7\\%) and SUN397 (72.0\\%). We\nrelease the code and models\nat~\\url{https://github.com/wanglimin/MRCNN-Scene-Recognition}.","url_abs":"http://arxiv.org/abs/1610.01119v2","url_pdf":"http://arxiv.org/pdf/1610.01119v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"knowledge-guided-disambiguation-for-large","repo_url":"https://github.com/wanglimin/MRCNN-Scene-Recognition","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"knowledge-guided-disambiguation-for-large","repo_url":"https://github.com/yjxiong/caffe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"scene-classification","task_name":"Scene Classification"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1610.01119","atlas_url":"https://app.syntology.ai/?focus=1610.01119","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}