Browse State-of-the-Art › Layout-to-Image Generation
Layout-to-Image Generation
24 papers with code · 11 benchmarks · 8 datasets archive 2025-07-28
Layout-to-image generation its the task to generate a scene based on the given layout. The layout describes the location of the objects to be included in the output image. In this section, you can find state-of-the-art leaderboards for Layout-to-image generation.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
11 leaderboard tables shown for this task, 11 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 11 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
24 shown of 24 papers with code (41 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Dec 2021 41 repositories listed Syntology ran 19 of 28 samples · 9 unverified · 5 pointer-only (licence)By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond.
-
10 Feb 2023 12 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedControlNet locks the production-ready large diffusion models, and reuses their deep and robust encoding layers pretrained with billions of images as a strong backbone to learn a diverse set of conditional controls.
-
20 Aug 2019 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Despite remarkable recent progress on both unconditional and conditional image synthesis, it remains a long-standing problem to learn generative models that are capable of synthesizing realistic and sharp images from…
-
4 Apr 2018 4 repositories listed Syntology ran 0 of 21 samples · 21 unverifiedTo overcome this limitation we propose a method for generating images from scene graphs, enabling explicitly reasoning about objects and their relationships.
-
25 Mar 2020 3 repositories listedThis paper focuses on a recent emerged task, layout-to-image, to learn generative models that are capable of synthesizing photo-realistic images from spatial layout (i.
-
13 Apr 2023 2 repositories listedIn this paper, we propose LayoutBench, a diagnostic benchmark for layout-guided image generation that examines four categories of spatial control skills: number, position, size, and shape.
-
30 Mar 2023 2 repositories listed Syntology ran 5 of 19 samples · 14 unverifiedTo overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal…
-
16 Dec 2019 2 repositories listed Syntology ran 2 of 17 samples · 15 unverifiedGenerating realistic images of complex visual scenes becomes challenging when one wishes to control the structure of the generated images.
-
11 Sep 2019 2 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedWe introduce a method for the generation of images from an input scene graph.
-
3 Mar 2025 1 repository listedThese models follow a one-stage framework: Encouraging the model to focus the attention map of each concept on its corresponding region by defining attention map-based losses.
-
18 Oct 2024 1 repository listed Syntology ran 13 of 13 samples · 0 unverified · 13 pointer-only (licence)The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions.
-
7 Sep 2024 1 repository listed Syntology ran 10 of 13 samples · 3 unverified · 13 pointer-only (licence)Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing.
-
18 Jul 2024 1 repository listedRecent breakthroughs in text-to-image diffusion models have significantly advanced the generation of high-fidelity, photo-realistic images from textual descriptions.
-
16 Jan 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)Current L2I models either suffer from poor editability via text or weak alignment between the generated image and the input layout.
-
16 Oct 2023 1 repository listed Syntology ran 10 of 12 samples · 2 unverified · 12 pointer-only (licence)Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects.
-
5 May 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Diffusion models have the ability to generate high quality images by denoising pure Gaussian noise images.
-
25 Mar 2023 1 repository listed Syntology ran 3 of 6 samples · 3 unverifiedIn this work, we explore the freestyle capability of the model, i.
-
17 Jan 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedLarge-scale text-to-image diffusion models have made amazing advances.
-
2 Jun 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedCompared to existing CNN-based and Transformer-based generation models that entangled modeling on pixel-level&patch-level and object-level&patch-level respectively, the proposed focal attention predicts the current…
-
4 Mar 2022 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedIn particular, the stuff layouts can take amorphous shapes and fill up the missing regions left out by the instance layouts.
-
25 Mar 2021 1 repository listedIn this paper, we propose a method for attribute controlled image synthesis from layout which allows to specify the appearance of individual objects without affecting the rest of the image.
-
22 Mar 2021 1 repository listedWe argue that these are caused by the lack of context-aware object and stuff feature encoding in their generators, and location-sensitive appearance representation in their discriminators.
-
22 Jun 2020 1 repository listedIn particular, layout-to-image generation models have gained significant attention due to their capability to generate realistic complex images containing distinct objects.
-
7 Apr 2020 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedIn our work, we address the novel problem of image manipulation from scene graphs, in which a user can edit images by merely applying changes in the nodes or edges of a semantic graph that is generated from the image.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections