Browse State-of-the-Art › Scene Generation
Scene Generation
122 papers with code · 6 benchmarks · 9 datasets archive 2025-07-28
make to t shirt an Ad with a little bit of action
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| GoogleEarth (5 rows) | GaussianCity | GaussianCity: Generative Gaussian Splatting for Unbounded 3D City... | code | Syntology ran 7 of 10 samples · 3 unverified | Compare |
| AVD (3 rows) | GSN | Unconstrained Scene Generation with Locally Conditioned Radiance Fields | code | — | Compare |
| Replica (3 rows) | GSN | Unconstrained Scene Generation with Locally Conditioned Radiance Fields | code | — | Compare |
| VizDoom (3 rows) | GSN | Unconstrained Scene Generation with Locally Conditioned Radiance Fields | code | — | Compare |
| OSM (2 rows) | InfiniteGAN | InfinityGAN: Towards Infinite-Pixel Image Synthesis | code | Syntology ran 7 of 14 samples · 7 unverified | Compare |
| KITTI (1 row) | GaussianCity | GaussianCity: Generative Gaussian Splatting for Unbounded 3D City... | code | Syntology ran 7 of 10 samples · 3 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
9 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 122 papers with code (309 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Jul 2020 7 repositories listedWe present a conceptually simple but effective funnel activation for image recognition tasks, called Funnel activation (FReLU), that extends ReLU and PReLU to a 2D activation by adding a negligible overhead of spatial…
-
10 Nov 2021 4 repositories listedHowever, current simulators for Embodied AI (EAI) challenges only provide simulated indoor scenes with a limited number of layouts.
-
24 Nov 2024 3 repositories listedTo train our SceneVLM, we collect over 610, 000 images from various public indoor datasets and implement a scene data generation pipeline with a semi-automated technique to establish relationships and estimate distances…
-
2 Dec 2020 3 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 2 pointer-only (licence)We have witnessed rapid progress on 3D-aware image synthesis, leveraging recent advances in generative visual models and neural rendering.
-
11 Dec 2024 2 repositories listedWe represent each scene with ego, agent, and map tokens and formulate autonomous driving as a unified token generation problem.
-
11 Dec 2024 2 repositories listed Syntology ran 6 of 23 samples · 17 unverified · 12 pointer-only (licence)However, existing T2I models show decayed performance in compositional image generation involving multiple objects and intricate relationships.
-
27 Aug 2024 2 repositories listedDesigns and artworks are ubiquitous across various creative fields, requiring graphic design skills and dedicated software to create compositions that include many graphical elements, such as logos, icons, symbols, and…
-
8 Aug 2024 2 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedThe emergence of foundation models as the "brain" of EAI agents for high-level task planning has shown promising results.
-
20 Nov 2023 2 repositories listed Syntology ran 11 of 12 samples · 1 unverifiedWe introduce a framework, the Pyramid Discrete Diffusion model (PDD), which employs scale-varied diffusion models to progressively generate high-quality outdoor scenes.
-
20 Apr 2021 2 repositories listedMoreover, object representations are often inferred using RNNs which do not scale well to large images or iterative refinement which avoids imposing an unnatural ordering on objects in an image but requires the a priori…
-
1 Mar 2021 2 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedWe introduce the GANformer, a novel and efficient type of transformer, and explore it for the task of visual generative modeling.
-
17 Dec 2020 2 repositories listedIn contrast, we do not use any appearance information, and implicitly learn object relations using the self-attention mechanism of transformers.
-
23 Jan 2020 2 repositories listedOne particular requirement for such robots is that they are able to understand spatial relations and can place objects in accordance with the spatial relations expressed by their user.
-
27 Dec 2019 2 repositories listedTo tackle this issue, in this work we consider learning the scene generation in a local context, and correspondingly design a local class-specific generative network with semantic maps as a guidance, which separately…
-
16 Dec 2019 2 repositories listed Syntology ran 2 of 17 samples · 15 unverifiedGenerating realistic images of complex visual scenes becomes challenging when one wishes to control the structure of the generated images.
-
26 Nov 2019 2 repositories listedFor the former, we use an unconditional progressive segmentation generation network that captures the distribution of realistic semantic scene layouts.
-
11 Sep 2019 2 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedWe introduce a method for the generation of images from an input scene graph.
-
30 Jul 2019 2 repositories listedGenerative latent-variable models are emerging as promising tools in robotics and reinforcement learning.
-
25 Jul 2019 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper we propose a neural message passing approach to augment an input 3D indoor scene with new objects matching their surroundings.
-
24 Jul 2019 2 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedRecently there is an increasing interest in scene generation within the research community.
-
25 Sep 2018 2 repositories listedWe propose a new probabilistic programming language for the design and analysis of perception systems, especially those based on machine learning.
-
12 Jul 2025 1 repository listedForecasting the evolution of 3D scenes and generating unseen scenarios via occupancy-based world models offers substantial potential for addressing corner cases in autonomous driving systems.
-
26 Jun 2025 1 repository listed Syntology ran 0 of 10 samples · 10 unverifiedAchieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of…
-
20 Jun 2025 1 repository listed Syntology ran 0 of 9 samples · 9 unverifiedPrior models and benchmarks focus on closed-loop motion simulation for initial agents in a scene.
-
8 May 2025 1 repository listedRecent advances in deep generative models (e.
-
7 May 2025 1 repository listedWe introduce a novel MCTS-based inference-time search strategy for diffusion models, enforce feasibility via projection and simulation, and release a dataset of over 44 million SE(3) scenes spanning five diverse…
-
3 May 2025 1 repository listedHowever, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control, fall short in capturing the complexity of the scene and integrating multi-modal information.
-
30 Apr 2025 1 repository listed Syntology ran 2 of 8 samples · 6 unverifiedTo address this issue, we propose HoloTime, a framework that integrates video diffusion models to generate panoramic videos from a single prompt or reference image, along with a 360-degree 4D scene reconstruction method…
-
7 Apr 2025 1 repository listedWe benchmark EP-Diffuser against two SotA models in terms of accuracy and plausibility of predictions on the Argoverse 2 dataset.
-
26 Mar 2025 1 repository listedWe present Free4D, a novel tuning-free framework for 4D scene generation from a single image.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections