{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semantic-aware-scene-recognition","title":"Semantic-Aware Scene Recognition","arxiv_id":"1909.02410","date":"2019-09-05","proceeding":null,"authors":["Alejandro López-Cifuentes","Marcos Escudero-Viñolo","Jesús Bescós","Álvaro García-Martín"],"abstract":"Scene recognition is currently one of the top-challenging research fields in computer vision. This may be due to the ambiguity between classes: images of several scene classes may share similar objects, which causes confusion among them. The problem is aggravated when images of a particular scene class are notably different. Convolutional Neural Networks (CNNs) have significantly boosted performance in scene recognition, albeit it is still far below from other recognition tasks (e.g., object or image recognition). In this paper, we describe a novel approach for scene recognition based on an end-to-end multi-modal CNN that combines image and context information by means of an attention module. Context information, in the shape of semantic segmentation, is used to gate features extracted from the RGB image by leveraging on information encoded in the semantic representation: the set of scene objects and stuff, and their relative locations. This gating process reinforces the learning of indicative scene content and enhances scene disambiguation by refocusing the receptive fields of the CNN towards them. Experimental results on four publicly available datasets show that the proposed approach outperforms every other state-of-the-art method while significantly reducing the number of network parameters. All the code and data used along this paper is available at https://github.com/vpulab/Semantic-Aware-Scene-Recognition","url_abs":"https://arxiv.org/abs/1909.02410v3","url_pdf":"https://arxiv.org/pdf/1909.02410v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"semantic-aware-scene-recognition","repo_url":"https://github.com/vpulab/Semantic-Aware-Scene-Recognition","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"scene-classification","task_name":"Scene Classification"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"places365","name":"Places365","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-recognition-on-ade20k","task":"Scene Recognition","dataset":"ADE20K","model":"Semantic-Aware Scene Recogniton (ResNet-18)","rank_in_archive_order":1,"of":1,"metrics":{"Top 1 Accuracy":"62.55"},"uses_additional_data":false},{"leaderboard":"/sota/scene-recognition-on-mit-indoors-scenes","task":"Scene Recognition","dataset":"MIT Indoor Scenes","model":"Semantic-Aware Scene Recognition (ResNet-50)","rank_in_archive_order":2,"of":3,"metrics":{"Accuracy":"87.10"},"uses_additional_data":false},{"leaderboard":"/sota/scene-recognition-on-places365","task":"Scene Recognition","dataset":"Places365","model":"Semantic-Aware Scene Recognition (ResNet-18)","rank_in_archive_order":2,"of":2,"metrics":{"Top 1 Accuracy":"56.51","Top 5 Accuracy":"86.00"},"uses_additional_data":false},{"leaderboard":"/sota/scene-recognition-on-sun397","task":"Scene Recognition","dataset":"SUN397","model":"Semantic-Aware Scene Recognition (ResNet-50)","rank_in_archive_order":2,"of":2,"metrics":{"Accuracy":"74.04"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1909.02410","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1909.02410"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vpulab/Semantic-Aware-Scene-Recognition","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"57c5a1a96edcb887","entry":"accuracy","repo":"vpulab/Semantic-Aware-Scene-Recognition","repo_kind":"official","path":"Libs/Utils/utils.py","file_url":"https://github.com/vpulab/Semantic-Aware-Scene-Recognition/blob/HEAD/Libs/Utils/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"57c5a1a96edcb887"}},{"code_sha256_prefix":"0516387258712998","entry":"getclassAccuracy","repo":"vpulab/Semantic-Aware-Scene-Recognition","repo_kind":"official","path":"Libs/Utils/utils.py","file_url":"https://github.com/vpulab/Semantic-Aware-Scene-Recognition/blob/HEAD/Libs/Utils/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0516387258712998"}},{"code_sha256_prefix":"826b8d2ca3bc9c42","entry":"unNormalizeImage","repo":"vpulab/Semantic-Aware-Scene-Recognition","repo_kind":"official","path":"Libs/Utils/utils.py","file_url":"https://github.com/vpulab/Semantic-Aware-Scene-Recognition/blob/HEAD/Libs/Utils/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"826b8d2ca3bc9c42"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}