{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/chatpainter-improving-text-to-image","title":"ChatPainter: Improving Text to Image Generation using Dialogue","arxiv_id":"1802.08216","date":"2018-02-22","proceeding":null,"authors":["Shikhar Sharma","Dendi Suhubdy","Vincent Michalski","Samira Ebrahimi Kahou","Yoshua Bengio"],"abstract":"Synthesizing realistic images from text descriptions on a dataset like\nMicrosoft Common Objects in Context (MS COCO), where each image can contain\nseveral objects, is a challenging task. Prior work has used text captions to\ngenerate images. However, captions might not be informative enough to capture\nthe entire image and insufficient for the model to be able to understand which\nobjects in the images correspond to which words in the captions. We show that\nadding a dialogue that further describes the scene leads to significant\nimprovement in the inception score and in the quality of generated images on\nthe MS COCO dataset.","url_abs":"http://arxiv.org/abs/1802.08216v1","url_pdf":"http://arxiv.org/pdf/1802.08216v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"text-to-image-generation-1","task_name":"Text to Image Generation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-image-generation-on-coco","task":"Text-to-Image Generation","dataset":"COCO (Common Objects in Context)","model":"ChatPainter","rank_in_archive_order":69,"of":69,"metrics":{"Inception score":"9.74"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.08216","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}