{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generating-multiple-objects-at-spatially","title":"Generating Multiple Objects at Spatially Distinct Locations","arxiv_id":"1901.00686","date":"2019-01-03","proceeding":"ICLR 2019 5","authors":["Tobias Hinz","Stefan Heinrich","Stefan Wermter"],"abstract":"Recent improvements to Generative Adversarial Networks (GANs) have made it\npossible to generate realistic images in high resolution based on natural\nlanguage descriptions such as image captions. Furthermore, conditional GANs\nallow us to control the image generation process through labels or even natural\nlanguage descriptions. However, fine-grained control of the image layout, i.e.\nwhere in the image specific objects should be located, is still difficult to\nachieve. This is especially true for images that should contain multiple\ndistinct objects at different spatial locations. We introduce a new approach\nwhich allows us to control the location of arbitrarily many objects within an\nimage by adding an object pathway to both the generator and the discriminator.\nOur approach does not need a detailed semantic layout but only bounding boxes\nand the respective labels of the desired objects are needed. The object pathway\nfocuses solely on the individual objects and is iteratively applied at the\nlocations specified by the bounding boxes. The global pathway focuses on the\nimage background and the general image layout. We perform experiments on the\nMulti-MNIST, CLEVR, and the more complex MS-COCO data set. Our experiments show\nthat through the use of the object pathway we can control object locations\nwithin images and can model complex scenes with multiple objects at various\nlocations. We further show that the object pathway focuses on the individual\nobjects and learns features relevant for these, while the global pathway\nfocuses on global image characteristics and the image background.","url_abs":"http://arxiv.org/abs/1901.00686v1","url_pdf":"http://arxiv.org/pdf/1901.00686v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generating-multiple-objects-at-spatially","repo_url":"https://github.com/tohinz/multiple-objects-gan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"conditional-image-generation","task_name":"Conditional Image Generation"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-image-generation-on-coco","task":"Text-to-Image Generation","dataset":"COCO (Common Objects in Context)","model":"AttnGAN + OP","rank_in_archive_order":61,"of":69,"metrics":{"FID":"33.35","Inception score":"24.76","SOA-C":"25.46"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-image-generation-on-coco","task":"Text-to-Image Generation","dataset":"COCO (Common Objects in Context)","model":"StackGAN + OP","rank_in_archive_order":65,"of":69,"metrics":{"FID":"55.30","Inception score":"12.12"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.00686","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}