{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/understanding-and-visualizing-deep-visual","title":"Understanding and Visualizing Deep Visual Saliency Models","arxiv_id":"1903.02501","date":"2019-03-06","proceeding":"CVPR 2019 6","authors":["Sen He","Hamed R. -Tavakoli","Ali Borji","Yang Mi","Nicolas Pugeault"],"abstract":"Recently, data-driven deep saliency models have achieved high performance and\nhave outperformed classical saliency models, as demonstrated by results on\ndatasets such as the MIT300 and SALICON. Yet, there remains a large gap between\nthe performance of these models and the inter-human baseline. Some outstanding\nquestions include what have these models learned, how and where they fail, and\nhow they can be improved. This article attempts to answer these questions by\nanalyzing the representations learned by individual neurons located at the\nintermediate layers of deep saliency models. To this end, we follow the steps\nof existing deep saliency models, that is borrowing a pre-trained model of\nobject recognition to encode the visual features and learning a decoder to\ninfer the saliency. We consider two cases when the encoder is used as a fixed\nfeature extractor and when it is fine-tuned, and compare the inner\nrepresentations of the network. To study how the learned representations depend\non the task, we fine-tune the same network using the same image set but for two\ndifferent tasks: saliency prediction versus scene classification. Our analyses\nreveal that: 1) some visual regions (e.g. head, text, symbol, vehicle) are\nalready encoded within various layers of the network pre-trained for object\nrecognition, 2) using modern datasets, we find that fine-tuning pre-trained\nmodels for saliency prediction makes them favor some categories (e.g. head)\nover some others (e.g. text), 3) although deep models of saliency outperform\nclassical models on natural images, the converse is true for synthetic stimuli\n(e.g. pop-out search arrays), an evidence of significant difference between\nhuman and data-driven saliency models, and 4) we confirm that, after-fine\ntuning, the change in inner-representations is mostly due to the task and not\nthe domain shift in the data.","url_abs":"http://arxiv.org/abs/1903.02501v3","url_pdf":"http://arxiv.org/pdf/1903.02501v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"understanding-and-visualizing-deep-visual","repo_url":"https://github.com/SenHe/uavdvsm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"},{"task_slug":"scene-classification","task_name":"Scene Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.02501","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1903.02501"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/SenHe/uavdvsm","reach":null}],"summary":{"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"82955b498f9b989d","entry":"cor","repo":"SenHe/uavdvsm","repo_kind":"official","path":"sal_train_pt.py","file_url":"https://github.com/SenHe/uavdvsm/blob/HEAD/sal_train_pt.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"82955b498f9b989d"}},{"code_sha256_prefix":"201779cda8e859d2","entry":"NSS","repo":"SenHe/uavdvsm","repo_kind":"official","path":"sal_train_pt.py","file_url":"https://github.com/SenHe/uavdvsm/blob/HEAD/sal_train_pt.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"201779cda8e859d2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}