{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/watch-what-you-just-said-image-captioning","title":"Watch What You Just Said: Image Captioning with Text-Conditional Attention","arxiv_id":"1606.04621","date":"2016-06-15","proceeding":null,"authors":["Luowei Zhou","Chenliang Xu","Parker Koch","Jason J. Corso"],"abstract":"Attention mechanisms have attracted considerable interest in image captioning\ndue to its powerful performance. However, existing methods use only visual\ncontent as attention and whether textual context can improve attention in image\ncaptioning remains unsolved. To explore this problem, we propose a novel\nattention mechanism, called \\textit{text-conditional attention}, which allows\nthe caption generator to focus on certain image features given previously\ngenerated text. To obtain text-related image features for our attention model,\nwe adopt the guiding Long Short-Term Memory (gLSTM) captioning architecture\nwith CNN fine-tuning. Our proposed method allows joint learning of the image\nembedding, text embedding, text-conditional attention and language model with\none network architecture in an end-to-end manner. We perform extensive\nexperiments on the MS-COCO dataset. The experimental results show that our\nmethod outperforms state-of-the-art captioning methods on various quantitative\nmetrics as well as in human evaluation, which supports the use of our\ntext-conditional attention in image captioning.","url_abs":"http://arxiv.org/abs/1606.04621v3","url_pdf":"http://arxiv.org/pdf/1606.04621v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"watch-what-you-just-said-image-captioning","repo_url":"https://github.com/LuoweiZhou/e2e-gLSTM-sc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}