{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/text-guided-attention-model-for-image","title":"Text-guided Attention Model for Image Captioning","arxiv_id":"1612.03557","date":"2016-12-12","proceeding":null,"authors":["Jonghwan Mun","Minsu Cho","Bohyung Han"],"abstract":"Visual attention plays an important role to understand images and\ndemonstrates its effectiveness in generating natural language descriptions of\nimages. On the other hand, recent studies show that language associated with an\nimage can steer visual attention in the scene during our cognitive process.\nInspired by this, we introduce a text-guided attention model for image\ncaptioning, which learns to drive visual attention using associated captions.\nFor this model, we propose an exemplar-based learning approach that retrieves\nfrom training data associated captions with each image, and use them to learn\nattention on visual features. Our attention model enables to describe a\ndetailed state of scenes by distinguishing small or confusable objects\neffectively. We validate our model on MS-COCO Captioning benchmark and achieve\nthe state-of-the-art performance in standard metrics.","url_abs":"http://arxiv.org/abs/1612.03557v1","url_pdf":"http://arxiv.org/pdf/1612.03557v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"text-guided-attention-model-for-image","repo_url":"https://github.com/vikramnitin9/nnfl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.03557","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}