{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-semi-supervised-framework-for-image","title":"A Semi-supervised Framework for Image Captioning","arxiv_id":"1611.05321","date":"2016-11-16","proceeding":null,"authors":["Wenhu Chen","Aurelien Lucchi","Thomas Hofmann"],"abstract":"State-of-the-art approaches for image captioning require supervised training\ndata consisting of captions with paired image data. These methods are typically\nunable to use unsupervised data such as textual data with no corresponding\nimages, which is a much more abundant commodity. We here propose a novel way of\nusing such textual data by artificially generating missing visual information.\nWe evaluate this learning approach on a newly designed model that detects\nvisual concepts present in an image and feed them to a reviewer-decoder\narchitecture with an attention mechanism. Unlike previous approaches that\nencode visual concepts using word embeddings, we instead suggest using regional\nimage features which capture more intrinsic information. The main benefit of\nthis architecture is that it synthesizes meaningful thought vectors that\ncapture salient image properties and then applies a soft attentive decoder to\ndecode the thought vectors and generate image captions. We evaluate our model\non both Microsoft COCO and Flickr30K datasets and demonstrate that this model\ncombined with our semi-supervised learning method can largely improve\nperformance and help the model to generate more accurate and diverse captions.","url_abs":"http://arxiv.org/abs/1611.05321v3","url_pdf":"http://arxiv.org/pdf/1611.05321v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-semi-supervised-framework-for-image","repo_url":"https://github.com/wenhuchen/ETHZ-Bootstrapped-Captioning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.05321","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}