{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/where-to-put-the-image-in-an-image-caption","title":"Where to put the Image in an Image Caption Generator","arxiv_id":"1703.09137","date":"2017-03-27","proceeding":null,"authors":["Marc Tanti","Albert Gatt","Kenneth P. Camilleri"],"abstract":"When a recurrent neural network language model is used for caption\ngeneration, the image information can be fed to the neural network either by\ndirectly incorporating it in the RNN -- conditioning the language model by\n`injecting' image features -- or in a layer following the RNN -- conditioning\nthe language model by `merging' image features. While both options are attested\nin the literature, there is as yet no systematic comparison between the two. In\nthis paper we empirically show that it is not especially detrimental to\nperformance whether one architecture is used or another. The merge architecture\ndoes have practical advantages, as conditioning by merging allows the RNN's\nhidden state vector to shrink in size by up to four times. Our results suggest\nthat the visual and linguistic modalities for caption generation need not be\njointly encoded by the RNN as that yields large, memory-intensive models with\nfew tangible advantages in performance; rather, the multimodal integration\nshould be delayed to a subsequent stage.","url_abs":"http://arxiv.org/abs/1703.09137v2","url_pdf":"http://arxiv.org/pdf/1703.09137v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/mtanti/where-image2","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/VinitSR7/Image-Caption-Generation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/bmy4415/DMLAB-intern","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/bprasannakumar/image_captioning_using_keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/jishubasak/Punny-Caption--Exploring-Image-Captioning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/kahotsang/image-captioning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/mrityunjayojha10/Image-Caption-Generation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/mtanti/rnn-role","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/nicolafan/image-captioning-cnn-rnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/shan1322/Neural-Style-Captioning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/simnyatsanga/image-caption-generator","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"where-to-put-the-image-in-an-image-caption","repo_url":"https://github.com/thomasvengal/deeplearning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"caption-generation","task_name":"Caption Generation"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}