{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-image-captioning-with-conditional","title":"Improving Image Captioning with Conditional Generative Adversarial Nets","arxiv_id":"1805.07112","date":"2018-05-18","proceeding":null,"authors":["Chen Chen","Shuai Mu","Wanpeng Xiao","Zexiong Ye","Liesi Wu","Qi Ju"],"abstract":"In this paper, we propose a novel\nconditional-generative-adversarial-nets-based image captioning framework as an\nextension of traditional reinforcement-learning (RL)-based encoder-decoder\narchitecture. To deal with the inconsistent evaluation problem among different\nobjective language metrics, we are motivated to design some \"discriminator\"\nnetworks to automatically and progressively determine whether generated caption\nis human described or machine generated. Two kinds of discriminator\narchitectures (CNN and RNN-based structures) are introduced since each has its\nown advantages. The proposed algorithm is generic so that it can enhance any\nexisting RL-based image captioning framework and we show that the conventional\nRL training method is just a special case of our approach. Empirically, we show\nconsistent improvements over all language evaluation metrics for different\nstate-of-the-art image captioning models. In addition, the well-trained\ndiscriminators can also be viewed as objective image captioning evaluators","url_abs":"http://arxiv.org/abs/1805.07112v4","url_pdf":"http://arxiv.org/pdf/1805.07112v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-image-captioning-with-conditional","repo_url":"https://github.com/Anjaney1999/image-captioning-seqgan","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.07112","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}