{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/paying-more-attention-to-saliency-image","title":"Paying More Attention to Saliency: Image Captioning with Saliency and Context Attention","arxiv_id":"1706.08474","date":"2017-06-26","proceeding":null,"authors":["Marcella Cornia","Lorenzo Baraldi","Giuseppe Serra","Rita Cucchiara"],"abstract":"Image captioning has been recently gaining a lot of attention thanks to the\nimpressive achievements shown by deep captioning architectures, which combine\nConvolutional Neural Networks to extract image representations, and Recurrent\nNeural Networks to generate the corresponding captions. At the same time, a\nsignificant research effort has been dedicated to the development of saliency\nprediction models, which can predict human eye fixations. Even though saliency\ninformation could be useful to condition an image captioning architecture, by\nproviding an indication of what is salient and what is not, research is still\nstruggling to incorporate these two techniques. In this work, we propose an\nimage captioning approach in which a generative recurrent neural network can\nfocus on different parts of the input image during the generation of the\ncaption, by exploiting the conditioning given by a saliency prediction model on\nwhich parts of the image are salient and which are contextual. We show, through\nextensive quantitative and qualitative experiments on large scale datasets,\nthat our model achieves superior performances with respect to captioning\nbaselines with and without saliency, and to different state of the art\napproaches combining saliency and captioning.","url_abs":"http://arxiv.org/abs/1706.08474v4","url_pdf":"http://arxiv.org/pdf/1706.08474v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-captioning-on-flickr30k-captions-test","task":"Image Captioning","dataset":"Flickr30k Captions test","model":"Cornia et al","rank_in_archive_order":2,"of":7,"metrics":{"BLEU-4":"21.3","CIDEr":"46.4","METEOR":"20.0","SPICE":"-"},"uses_additional_data":true}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1706.08474","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}