{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-diverse-and-natural-image","title":"Towards Diverse and Natural Image Descriptions via a Conditional GAN","arxiv_id":"1703.06029","date":"2017-03-17","proceeding":"ICCV 2017 10","authors":["Bo Dai","Sanja Fidler","Raquel Urtasun","Dahua Lin"],"abstract":"Despite the substantial progress in recent years, the image captioning\ntechniques are still far from being perfect.Sentences produced by existing\nmethods, e.g. those based on RNNs, are often overly rigid and lacking in\nvariability. This issue is related to a learning principle widely used in\npractice, that is, to maximize the likelihood of training samples. This\nprinciple encourages high resemblance to the \"ground-truth\" captions while\nsuppressing other reasonable descriptions. Conventional evaluation metrics,\ne.g. BLEU and METEOR, also favor such restrictive methods. In this paper, we\nexplore an alternative approach, with the aim to improve the naturalness and\ndiversity -- two essential properties of human expression. Specifically, we\npropose a new framework based on Conditional Generative Adversarial Networks\n(CGAN), which jointly learns a generator to produce descriptions conditioned on\nimages and an evaluator to assess how well a description fits the visual\ncontent. It is noteworthy that training a sequence generator is nontrivial. We\novercome the difficulty by Policy Gradient, a strategy stemming from\nReinforcement Learning, which allows the generator to receive early feedback\nalong the way. We tested our method on two large datasets, where it performed\ncompetitively against real people in our user study and outperformed other\nmethods on various tasks.","url_abs":"http://arxiv.org/abs/1703.06029v3","url_pdf":"http://arxiv.org/pdf/1703.06029v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-diverse-and-natural-image","repo_url":"https://github.com/doubledaibo/gancaption_iccv2017","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1703.06029","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}