Papers › Object-Centric Image Generation from Layouts

Object-Centric Image Generation from Layouts

16 Mar 2020arXiv:2003.07449archive 2025-07-28

Tristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm, Shikhar Sharma

Despite recent impressive results on single-object and single-domain image generation, the generation of complex scenes with multiple objects remains challenging. In this paper, we start with the idea that a model must be able to understand individual objects and relationships between objects in order to generate complex scenes well. Our layout-to-image-generation method, which we call Object-Centric Generative Adversarial Network (or OC-GAN), relies on a novel Scene-Graph Similarity Module (SGSM). The SGSM learns representations of the spatial relationships between objects in the scene, which lead to our model's improved layout-fidelity. We also propose changes to the conditioning mechanism of the generator that enhance its object instance-awareness. Apart from improving image quality, our contributions mitigate two failure modes in previous approaches: (1) spurious objects being generated without corresponding bounding boxes in the layout, and (2) overlapping bounding boxes in the layout leading to merged objects in images. Extensive quantitative evaluation and ablation studies demonstrate the impact of our contributions, with our model outperforming previous state-of-the-art approaches on both the COCO-Stuff and Visual Genome datasets. Finally, we address an important limitation of evaluation metrics used in previous works by introducing SceneFID -- an object-centric adaptation of the popular Fr{\'e}chet Inception Distance metric, that is better suited for multi-object images.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationLayout-to-Image GenerationObject

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Layout-to-Image Generation COCO-Stuff 128x128 OC-GAN FID 36.31 #4 of 5 Archive leaderboard report
Layout-to-Image Generation COCO-Stuff 128x128 OC-GAN Inception Score 14.6 #4 of 5 Archive leaderboard report
Layout-to-Image Generation COCO-Stuff 128x128 OC-GAN SceneFID 16.76 #4 of 5 Archive leaderboard report
Layout-to-Image Generation COCO-Stuff 256x256 OC-GAN FID 41.65 #3 of 5 Archive leaderboard report
Layout-to-Image Generation COCO-Stuff 256x256 OC-GAN Inception Score 17.8 #3 of 5 Archive leaderboard report
Layout-to-Image Generation COCO-Stuff 64x64 OC-GAN FID 29.57 #1 of 5 Archive leaderboard report
Layout-to-Image Generation COCO-Stuff 64x64 OC-GAN Inception Score 10.8 #1 of 5 Archive leaderboard report
Layout-to-Image Generation Visual Genome 128x128 OC-GAN FID 28.26 #3 of 5 Archive leaderboard report
Layout-to-Image Generation Visual Genome 128x128 OC-GAN Inception Score 12.3 #3 of 5 Archive leaderboard report
Layout-to-Image Generation Visual Genome 128x128 OC-GAN SceneFID 9.63 #3 of 5 Archive leaderboard report
Layout-to-Image Generation Visual Genome 256x256 OC-GAN FID 40.85 #4 of 4 Archive leaderboard report
Layout-to-Image Generation Visual Genome 256x256 OC-GAN Inception Score 14.7 #4 of 4 Archive leaderboard report
Layout-to-Image Generation Visual Genome 64x64 OC-GAN FID 20.27 #1 of 4 Archive leaderboard report
Layout-to-Image Generation Visual Genome 64x64 OC-GAN Inception Score 9.3 #1 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections