{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/guided-open-vocabulary-image-captioning-with","title":"Guided Open Vocabulary Image Captioning with Constrained Beam Search","arxiv_id":"1612.00576","date":"2016-12-02","proceeding":"EMNLP 2017 9","authors":["Peter Anderson","Basura Fernando","Mark Johnson","Stephen Gould"],"abstract":"Existing image captioning models do not generalize well to out-of-domain\nimages containing novel scenes or objects. This limitation severely hinders the\nuse of these models in real world applications dealing with images in the wild.\nWe address this problem using a flexible approach that enables existing deep\ncaptioning architectures to take advantage of image taggers at test time,\nwithout re-training. Our method uses constrained beam search to force the\ninclusion of selected tag words in the output, and fixed, pretrained word\nembeddings to facilitate vocabulary expansion to previously unseen tag words.\nUsing this approach we achieve state of the art results for out-of-domain\ncaptioning on MSCOCO (and improved results for in-domain captioning). Perhaps\nsurprisingly, our results significantly outperform approaches that incorporate\nthe same tag predictions into the learning algorithm. We also show that we can\nsignificantly improve the quality of generated ImageNet captions by leveraging\nground-truth labels.","url_abs":"http://arxiv.org/abs/1612.00576v2","url_pdf":"http://arxiv.org/pdf/1612.00576v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"guided-open-vocabulary-image-captioning-with","repo_url":"https://github.com/nocaps-org/updown-baseline","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"tag","task_name":"TAG"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.00576","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}