Papers › Generating Multiple Objects at Spatially Distinct Locations

Generating Multiple Objects at Spatially Distinct Locations

3 Jan 2019ICLR 2019 5arXiv:1901.00686archive 2025-07-28

Tobias Hinz, Stefan Heinrich, Stefan Wermter

Recent improvements to Generative Adversarial Networks (GANs) have made it possible to generate realistic images in high resolution based on natural language descriptions such as image captions. Furthermore, conditional GANs allow us to control the image generation process through labels or even natural language descriptions. However, fine-grained control of the image layout, i.e. where in the image specific objects should be located, is still difficult to achieve. This is especially true for images that should contain multiple distinct objects at different spatial locations. We introduce a new approach which allows us to control the location of arbitrarily many objects within an image by adding an object pathway to both the generator and the discriminator. Our approach does not need a detailed semantic layout but only bounding boxes and the respective labels of the desired objects are needed. The object pathway focuses solely on the individual objects and is iteratively applied at the locations specified by the bounding boxes. The global pathway focuses on the image background and the general image layout. We perform experiments on the Multi-MNIST, CLEVR, and the more complex MS-COCO data set. Our experiments show that through the use of the object pathway we can control object locations within images and can model complex scenes with multiple objects at various locations. We further show that the object pathway focuses on the individual objects and learns features relevant for these, while the global pathway focuses on global image characteristics and the image background.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tohinz/multiple-objects-gan officialmentioned in paperpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Conditional Image GenerationImage GenerationObjectText-to-Image Generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-Image Generation COCO (Common Objects in Context) AttnGAN + OP FID 33.35 #61 of 69 Archive leaderboard report
Text-to-Image Generation COCO (Common Objects in Context) AttnGAN + OP Inception score 24.76 #61 of 69 Archive leaderboard report
Text-to-Image Generation COCO (Common Objects in Context) AttnGAN + OP SOA-C 25.46 #61 of 69 Archive leaderboard report
Text-to-Image Generation COCO (Common Objects in Context) StackGAN + OP FID 55.30 #65 of 69 Archive leaderboard report
Text-to-Image Generation COCO (Common Objects in Context) StackGAN + OP Inception score 12.12 #65 of 69 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections