{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/top-down-visual-saliency-guided-by-captions","title":"Top-down Visual Saliency Guided by Captions","arxiv_id":"1612.07360","date":"2016-12-21","proceeding":"CVPR 2017 7","authors":["Vasili Ramanishka","Abir Das","Jianming Zhang","Kate Saenko"],"abstract":"Neural image/video captioning models can generate accurate descriptions, but\ntheir internal process of mapping regions to words is a black box and therefore\ndifficult to explain. Top-down neural saliency methods can find important\nregions given a high-level semantic task such as object classification, but\ncannot use a natural language sentence as the top-down input for the task. In\nthis paper, we propose Caption-Guided Visual Saliency to expose the\nregion-to-word mapping in modern encoder-decoder networks and demonstrate that\nit is learned implicitly from caption training data, without any pixel-level\nannotations. Our approach can produce spatial or spatiotemporal heatmaps for\nboth predicted captions, and for arbitrary query sentences. It recovers\nsaliency without the overhead of introducing explicit attention layers, and can\nbe used to analyze a variety of existing model architectures and improve their\ndesign. Evaluation on large-scale video and image datasets demonstrates that\nour approach achieves comparable captioning performance with existing methods\nwhile providing more accurate saliency heatmaps. Our code is available at\nvisionlearninggroup.github.io/caption-guided-saliency/.","url_abs":"http://arxiv.org/abs/1612.07360v2","url_pdf":"http://arxiv.org/pdf/1612.07360v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"top-down-visual-saliency-guided-by-captions","repo_url":"https://github.com/IgnacioHeredia/plant_classification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"top-down-visual-saliency-guided-by-captions","repo_url":"https://github.com/indigo-dc/conus-classification-theano","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"top-down-visual-saliency-guided-by-captions","repo_url":"https://github.com/indigo-dc/phytoplankton-classification-theano","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"top-down-visual-saliency-guided-by-captions","repo_url":"https://github.com/indigo-dc/seeds-classification-theano","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"top-down-visual-saliency-guided-by-captions","repo_url":"https://github.com/laramaktub/plant-classification-tf-train","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"top-down-visual-saliency-guided-by-captions","repo_url":"https://github.com/sudatta0993/Dynamic-Congestion-Prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-captioning","task_name":"Video Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}