{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/paying-attention-to-descriptions-generated-by","title":"Paying Attention to Descriptions Generated by Image Captioning Models","arxiv_id":"1704.07434","date":"2017-04-24","proceeding":"ICCV 2017 10","authors":["Hamed R. -Tavakoli","Rakshith Shetty","Ali Borji","Jorma Laaksonen"],"abstract":"To bridge the gap between humans and machines in image understanding and\ndescribing, we need further insight into how people describe a perceived scene.\nIn this paper, we study the agreement between bottom-up saliency-based visual\nattention and object referrals in scene description constructs. We investigate\nthe properties of human-written descriptions and machine-generated ones. We\nthen propose a saliency-boosted image captioning model in order to investigate\nbenefits from low-level cues in language models. We learn that (1) humans\nmention more salient objects earlier than less salient ones in their\ndescriptions, (2) the better a captioning model performs, the better attention\nagreement it has with human descriptions, (3) the proposed saliency-boosted\nmodel, compared to its baseline form, does not improve significantly on the MS\nCOCO database, indicating explicit bottom-up boosting does not help when the\ntask is well learnt and tuned on a data, (4) a better generalization is,\nhowever, observed for the saliency-boosted model on unseen data.","url_abs":"http://arxiv.org/abs/1704.07434v3","url_pdf":"http://arxiv.org/pdf/1704.07434v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"paying-attention-to-descriptions-generated-by","repo_url":"https://github.com/HemanthTejaY/Deep-Learning-Image-Captioning---A-comparitive-study","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"paying-attention-to-descriptions-generated-by","repo_url":"https://github.com/rakshithShetty/captionGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1704.07434","atlas_url":"https://app.syntology.ai/?focus=1704.07434","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}