{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-what-and-where-to-draw","title":"Learning What and Where to Draw","arxiv_id":"1610.02454","date":"2016-10-08","proceeding":"NeurIPS 2016 12","authors":["Scott Reed","Zeynep Akata","Santosh Mohan","Samuel Tenka","Bernt Schiele","Honglak Lee"],"abstract":"Generative Adversarial Networks (GANs) have recently demonstrated the\ncapability to synthesize compelling real-world images, such as room interiors,\nalbum covers, manga, faces, birds, and flowers. While existing models can\nsynthesize images based on global constraints such as a class label or caption,\nthey do not provide control over pose or object location. We propose a new\nmodel, the Generative Adversarial What-Where Network (GAWWN), that synthesizes\nimages given instructions describing what content to draw in which location. We\nshow high-quality 128 x 128 image synthesis on the Caltech-UCSD Birds dataset,\nconditioned on both informal text descriptions and also object location. Our\nsystem exposes control over both the bounding box around the bird and its\nconstituent parts. By modeling the conditional distributions over part\nlocations, our system also enables conditioning on arbitrary subsets of parts\n(e.g. only the beak and tail), yielding an efficient interface for picking part\nlocations. We also show preliminary results on the more challenging domain of\ntext- and location-controllable synthesis of images of human actions on the\nMPII Human Pose dataset.","url_abs":"http://arxiv.org/abs/1610.02454v1","url_pdf":"http://arxiv.org/pdf/1610.02454v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-image-generation-on-cub","task":"Text-to-Image Generation","dataset":"CUB","model":"GAWWN","rank_in_archive_order":14,"of":20,"metrics":{"FID":"67.22","Inception score":"3.62"},"uses_additional_data":true}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1610.02454","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}