Browse State-of-the-Art › Image Generation

Image Generation

3,102 papers with code · 93 benchmarks · 79 datasets archive 2025-07-28

Computer VisionMedicalMiscellaneousNatural Language Processing

Image Generation (synthesis) is the task of generating new images from an existing dataset.

  • Unconditional generation refers to generating samples unconditionally from the dataset, i.e. p(y)
  • Conditional image generation (subtask) refers to generating samples conditionally from the dataset, based on a label, i.e. p(y|x).

In this section, you can find state-of-the-art leaderboards for unconditional generation. For conditional generation, and other types of image generations, refer to the subtasks.

( Image credit: StyleGAN )

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

93 leaderboard tables shown for this task, 93 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 93 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
ImageNet 256x256 (94 rows) SiT-XL/2 + UCGM-S (E2E-VAE + 40 sampling steps + CFG) Unified Continuous Generative Models code Syntology ran 16 of 21 samples · 5 unverified Compare
CIFAR-10 (78 rows) GMem Generative Modeling with Explicit Memory code Syntology ran 8 of 15 samples · 7 unverified Compare
ImageNet 64x64 (65 rows) SIMS Self-Improving Diffusion Models with Synthetic Data — — Compare
ImageNet 512x512 (52 rows) EDM2-L + DDO (SD-VAE, 25 steps, DPM-Solver-v3) Direct Discriminative Optimization: Your Likelihood-Based Visual... code — Compare
FFHQ 256 x 256 (51 rows) StyleSAN-XL SAN: Inducing Metrizability of GAN with Discriminative Normalized... code Syntology ran 2 of 2 samples · 0 unverified Compare
CelebA 64x64 (39 rows) DDPM-IP Input Perturbation Reduces Exposure Bias in Diffusion Models code Syntology ran 2 of 3 samples · 1 unverified Compare
ImageNet 32x32 (35 rows) PaGoDA PaGoDA: Progressive Growing of a One-Step Generator from a... code Syntology ran 15 of 17 samples · 2 unverified Compare
LSUN Bedroom 256 x 256 (32 rows) Diffusion ProjectedGAN Diffusion-GAN: Training GANs with Diffusion code Syntology ran 13 of 18 samples · 5 unverified Compare
STL-10 (31 rows) Diffusion ProjectedGAN Diffusion-GAN: Training GANs with Diffusion code Syntology ran 13 of 18 samples · 5 unverified Compare
LSUN Churches 256 x 256 (27 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
ImageNet 128x128 (23 rows) SiD2 Simpler Diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion — — Compare
FFHQ 1024 x 1024 (20 rows) StyleSAN-XL SAN: Inducing Metrizability of GAN with Discriminative Normalized... code Syntology ran 2 of 2 samples · 0 unverified Compare
CelebA-HQ 256x256 (19 rows) RDM Relay Diffusion: Unifying diffusion process across resolutions for... code Syntology ran 12 of 17 samples · 5 unverified Compare
CelebA 256x256 (17 rows) Efficient-VDVAE Efficient-VDVAE: Less is more code — Compare
MNIST (15 rows) Locally Masked PixelCNN (8 orders) Locally Masked Convolution for Autoregressive Models code — Compare
WISE (14 rows) MindOmni (w/ cot) MindOmni: Unleashing Reasoning Generation in Vision Language... code — Compare
FFHQ-U (13 rows) Alias-Free-R Alias-Free Generative Adversarial Networks code Syntology ran 18 of 31 samples · 13 unverified Compare
FFHQ (12 rows) Anycost GAN Anycost GANs for Interactive Image Synthesis and Editing code Syntology ran 0 of 1 samples · 1 unverified Compare
Binarized MNIST (10 rows) CR-NVAE Consistency Regularization for Variational Auto-Encoders code Syntology ran 0 of 3 samples · 3 unverified Compare
CelebA-HQ 1024x1024 (10 rows) StyleSwin StyleSwin: Transformer-based GAN for High-resolution Image Generation code Syntology ran 5 of 5 samples · 0 unverified Compare
CIFAR-100 (9 rows) LeCAM (StyleGAN2 + ADA) Regularizing Generative Adversarial Networks under Limited Data code Syntology ran 3 of 9 samples · 6 unverified Compare
AFHQ Cat (8 rows) Vision-aided GAN Ensembling Off-the-shelf Models for GAN Training code Syntology ran 9 of 15 samples · 6 unverified Compare
LSUN Cat 256 x 256 (8 rows) Vision-aided GAN Ensembling Off-the-shelf Models for GAN Training code Syntology ran 9 of 15 samples · 6 unverified Compare
AFHQV2 (7 rows) Polarity-StyleGAN3 Polarity Sampling: Quality and Diversity Control of Pre-Trained... code — Compare
CelebA-HQ 128x128 (7 rows) U-Net GAN A U-Net Based Discriminator for Generative Adversarial Networks code — Compare
Fashion-MNIST (7 rows) GLF+perceptual loss (ours) Generative Latent Flow code Syntology ran 4 of 5 samples · 1 unverified Compare
TextAtlasEval (7 rows) Grok3 — — — Compare
AFHQ Dog (6 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
Cityscapes (6 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
CLEVR (6 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
LSUN Horse 256 x 256 (6 rows) Vision-aided GAN Ensembling Off-the-shelf Models for GAN Training code Syntology ran 9 of 15 samples · 6 unverified Compare
AFHQ Wild (5 rows) Vision-aided GAN Ensembling Off-the-shelf Models for GAN Training code Syntology ran 9 of 15 samples · 6 unverified Compare
CelebA 128x128 (5 rows) U-Net GAN A U-Net Based Discriminator for Generative Adversarial Networks code — Compare
Places50 (5 rows) SinDiffusion SinDiffusion: Learning a Diffusion Model from a Single Natural Image code Syntology ran 13 of 20 samples · 7 unverified Compare
ARKitScenes (4 rows) GAUDI GAUDI: A Neural Architect for Immersive 3D Scene Generation code — Compare
CUB 128 x 128 (4 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
Pokemon 256x256 (4 rows) StyleGAN-XL StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets code Syntology ran 14 of 19 samples · 5 unverified Compare
Replica (4 rows) GAUDI GAUDI: A Neural Architect for Immersive 3D Scene Generation code — Compare
Stanford Cars (4 rows) Projected GANs Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
Stanford Dogs (4 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
VizDoom (4 rows) GAUDI GAUDI: A Neural Architect for Immersive 3D Scene Generation code — Compare
VLN-CE (4 rows) GAUDI GAUDI: A Neural Architect for Immersive 3D Scene Generation code — Compare
ADE-Indoor (3 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
CAT 256x256 (3 rows) StyleGAN2 + DA + RLC (Ours) Regularizing Generative Adversarial Networks under Limited Data code Syntology ran 3 of 9 samples · 6 unverified Compare
CelebA-HQ 64x64 (3 rows) COCO-GAN COCO-GAN: Generation by Parts via Conditional Coordinating code Syntology ran 0 of 6 samples · 6 unverified Compare
CIFAR-10 (10% data) (3 rows) DiffAugment-StyleGAN2 Differentiable Augmentation for Data-Efficient GAN Training code Syntology ran 36 of 58 samples · 22 unverified Compare
CIFAR-10 (20% data) (3 rows) DiffAugment-StyleGAN2 Differentiable Augmentation for Data-Efficient GAN Training code Syntology ran 36 of 58 samples · 22 unverified Compare
FFHQ 128 x 128 (3 rows) DDPM-IP Input Perturbation Reduces Exposure Bias in Diffusion Models code Syntology ran 2 of 3 samples · 1 unverified Compare
FFHQ 512 x 512 (3 rows) StyleSAN-XL SAN: Inducing Metrizability of GAN with Discriminative Normalized... code Syntology ran 2 of 2 samples · 0 unverified Compare
LSUN Bedroom (3 rows) StyleGAN A Style-Based Generator Architecture for Generative Adversarial Networks code Syntology ran 226 of 396 samples · 170 unverified Compare
LSUN Bedroom 64 x 64 (3 rows) WGAN-GP + TTUR + Alex-Adam Fundamental Benefit of Alternating Updates in Minimax Optimization code Syntology ran 1 of 2 samples · 1 unverified Compare
MetFaces (3 rows) t-Stylegan3-ada (NVIDIA pre-trained) Signature and Log-signature for the Study of Empirical... code — Compare
MetFaces-U (3 rows) Alias-Free-R Alias-Free Generative Adversarial Networks code Syntology ran 18 of 31 samples · 13 unverified Compare
ObjectsRoom (3 rows) GENESIS-V2 GENESIS-V2: Inferring Unordered Object Representations without... code — Compare
Pokemon 1024x1024 (3 rows) StyleGAN-XL StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets code Syntology ran 14 of 19 samples · 5 unverified Compare
ShapeStacks (3 rows) GENESIS-V2 GENESIS-V2: Inferring Unordered Object Representations without... code — Compare
Stacked MNIST (3 rows) VAEBM VAEBM: A Symbiosis between Variational Autoencoders and Energy-based Models code Syntology ran 6 of 9 samples · 3 unverified Compare
TiO_2 nanoparticle (3 rows) F-ANcGAN F-ANcGAN: An Attention-Enhanced Cycle Consistent Generative... code — Compare
AFHQ-v2 64x64 (2 rows) SiDA-EDM Adversarial Score identity Distillation: Rapidly Surpassing the... code Syntology ran 2 of 3 samples · 1 unverified Compare
FFHQ 64x64 (2 rows) SiDA-EDM Adversarial Score identity Distillation: Rapidly Surpassing the... code Syntology ran 2 of 3 samples · 1 unverified Compare
iNaturalist 2019 (2 rows) StyeGAN2 + NoisyTwins NoisyTwins: Class-Consistent and Diverse Image Generation through StyleGANs code — Compare
LSUN Bedroom 128 x 128 (2 rows) LadaGAN Efficient generative adversarial networks using linear... code — Compare
LSUN Car 512 x 384 (2 rows) Polarity-StyleGAN2 Polarity Sampling: Quality and Diversity Control of Pre-Trained... code — Compare
Oxford 102 Flowers 256 x 256 (2 rows) Projected GAN Projected GANs Converge Faster code Syntology ran 38 of 49 samples · 11 unverified Compare
RC-49 (2 rows) cDR-RS Efficient Subsampling of Realistic Images From GANs Conditional on... code — Compare
1,078 People 3D Faces Collection Data (1 row) Sessiz çığlık 0-Hecke modules for row-strict dual immaculate functions — — Compare
25% ImageNet 128x128 (1 row) LeCAM + DA Regularizing Generative Adversarial Networks under Limited Data code Syntology ran 3 of 9 samples · 6 unverified Compare
‘ไอซ์ ปรีชญา’ ลืมปิดไลฟ์สดตอนอาบน้ำ คลิปถูกคนดีแชร์ออนไลน์ (1 row) as Image Generation From Small Datasets via Batch Statistics Adaptation code Syntology ran 3 of 7 samples · 4 unverified Compare
CelebA (1 row) FInCFlow FInC Flow: Fast and Invertible k ×k Convolutions for Normalizing Flows — — Compare
CelebA-HQ (1 row) DDPM Normalizing Flow-Based Metric for Image Generation code — Compare
CelebA-HQ 512x512 (1 row) WaveDiff Wavelet Diffusion Models are fast and scalable Image Generators code — Compare
Cityscapes-25K 256x512 (1 row) SB-GAN Semantic Bottleneck Scene Generation code — Compare
Cityscapes-5K 256x512 (1 row) SB-GAN Semantic Bottleneck Scene Generation code — Compare
EMNIST-Letters (1 row) Spiking-Diffusion Spiking-Diffusion: Vector Quantized Discrete Diffusion Model with... code — Compare
FFHQ 64x64 - 4x upscaling (1 row) PFGM++ PFGM++: Unlocking the Potential of Physics-Inspired Generative Models code Syntology ran 2 of 3 samples · 1 unverified Compare
GQN (1 row) GENESIS GENESIS: Generative Scene Inference and Sampling with... code — Compare
ImageNet 256x256 - 1 labeled data per class (1 row) DPT — — — Compare
ImageNet 256x256 - 1% labeled data (1 row) DPT — — — Compare
ImageNet 256x256 - 2 labeled data per class (1 row) DPT — — — Compare
ImageNet 256x256 - 5 labeled data per class (1 row) DPT — — — Compare
Indian Celebs 256 x 256 (1 row) MSG-StyleGAN MSG-GAN: Multi-Scale Gradients for Generative Adversarial Networks code Syntology ran 3 of 15 samples · 12 unverified Compare
KMNIST (1 row) Spiking-Diffusion Spiking-Diffusion: Vector Quantized Discrete Diffusion Model with... code — Compare
Landscapes 256 x 256 (1 row) CIPS Image Generators with Conditionally-Independent Pixel Synthesis code — Compare
LLVIP (1 row) pix2pix LLVIP: A Visible-infrared Paired Dataset for Low-light Vision code — Compare
LSUN (1 row) BigGAN + gSR Improving GANs for Long-Tailed Data through Group Spectral Regularization code Syntology ran 0 of 3 samples · 3 unverified Compare
LSUN Car 256 x 256 (1 row) StyleGAN2 Analyzing and Improving the Image Quality of StyleGAN code Syntology ran 29 of 92 samples · 63 unverified Compare
LSUN tower 64x64 (1 row) DDPM-IP Input Perturbation Reduces Exposure Bias in Diffusion Models code Syntology ran 2 of 3 samples · 1 unverified Compare
Multi-dSprites (1 row) GENESIS GENESIS: Generative Scene Inference and Sampling with... code — Compare
NASA Perseverance (1 row) Stylegan2-ada Signature and Log-signature for the Study of Empirical... code — Compare
Oxford 102 Flowers 128x128 (1 row) QSNGAN Quaternion Generative Adversarial Networks code — Compare
Satellite-Buildings 256 x 256 (1 row) CIPS Image Generators with Conditionally-Independent Pixel Synthesis code — Compare
Satellite-Landscapes 256 x 256 (1 row) CIPS Image Generators with Conditionally-Independent Pixel Synthesis code — Compare
SDSS Galaxies (1 row) AstroDDPM Realistic galaxy image simulation via score-based generative models code Syntology ran 5 of 6 samples · 1 unverified Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

79 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 79 until expanded.

Subtasks archive 2025-07-28

23 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 3,102 papers with code (6,689 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 3 Dec 2019 126 repositories listed Syntology ran 29 of 92 samples · 63 unverified · 18 pointer-only (licence)
    Overall, our improved model redefines the state of the art in unconditional image modeling, both in terms of existing distribution quality metrics as well as perceived image quality.
  • 26 Jan 2017 120 repositories listed Syntology ran 13 of 22 samples · 9 unverified · 13 pointer-only (licence)
    We introduce a new algorithm named WGAN, an alternative to traditional GAN training.
  • 27 Oct 2017 115 repositories listed Syntology ran 24 of 89 samples · 65 unverified · 19 pointer-only (licence)
    We describe a new training methodology for generative adversarial networks.
  • 31 Mar 2017 110 repositories listed Syntology ran 27 of 49 samples · 22 unverified · 20 pointer-only (licence)
    Generative Adversarial Networks (GANs) are powerful generative models, but suffer from training instability.
  • 12 Dec 2018 83 repositories listed Syntology ran 226 of 396 samples · 170 unverified · 229 pointer-only (licence)
    We propose an alternative generator architecture for generative adversarial networks, borrowing from style transfer literature.
  • 26 Jun 2017 71 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 1 pointer-only (licence)
    Generative Adversarial Networks (GANs) excel at creating realistic images with complex models for which maximum likelihood is infeasible.
  • 19 Jun 2020 70 repositories listed Syntology ran 178 of 253 samples · 75 unverified · 62 pointer-only (licence)
    We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics.
  • 21 May 2018 48 repositories listed Syntology ran 18 of 81 samples · 63 unverified · 1 pointer-only (licence)
    In this paper, we propose the Self-Attention Generative Adversarial Network (SAGAN) which allows attention-driven, long-range dependency modeling for image generation tasks.
  • 2 May 2019 47 repositories listed Syntology ran 2 of 11 samples · 9 unverified · 3 pointer-only (licence)
    We introduce SinGAN, an unconditional generative model that can be learned from a single natural image.
  • 10 Jun 2016 46 repositories listed Syntology ran 1 of 2 samples · 1 unverified
    We present a variety of new architectural features and training procedures that we apply to the generative adversarial networks (GANs) framework.
  • 20 Dec 2021 41 repositories listed Syntology ran 19 of 28 samples · 9 unverified · 5 pointer-only (licence)
    By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond.
  • 17 May 2016 39 repositories listed Syntology ran 9 of 19 samples · 10 unverified · 6 pointer-only (licence)
    Automatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal.
  • 16 Feb 2018 38 repositories listed Syntology ran 16 of 31 samples · 15 unverified · 14 pointer-only (licence)
    One of the challenges in the study of generative adversarial networks is the instability of its training.
  • 12 Jun 2016 38 repositories listed Syntology ran 5 of 6 samples · 1 unverified
    This paper describes InfoGAN, an information-theoretic extension to the Generative Adversarial Network that is able to learn disentangled representations in a completely unsupervised manner.
  • 30 Oct 2016 37 repositories listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)
    We expand on previous work for image quality assessment to provide two new analyses for assessing the discriminability and diversity of samples from class-conditional image synthesis models.
  • 28 Sep 2018 35 repositories listed Syntology ran 14 of 41 samples · 27 unverified · 4 pointer-only (licence)
    Despite recent progress in generative image modeling, successfully generating high-resolution, diverse samples from complex datasets such as ImageNet remains an elusive goal.
  • 27 May 2016 35 repositories listed Syntology ran 43 of 70 samples · 27 unverified · 33 pointer-only (licence)
    Unsupervised learning of probabilistic models is a central yet challenging problem in machine learning.
  • 6 Oct 2020 29 repositories listed Syntology ran 28 of 50 samples · 22 unverified · 5 pointer-only (licence)
    Denoising diffusion probabilistic models (DDPMs) have achieved high quality image generation without adversarial training, yet they require simulating a Markov chain for many steps to produce a sample.
  • 11 Jun 2020 28 repositories listed Syntology ran 5 of 29 samples · 24 unverified · 4 pointer-only (licence)
    We also find that the widely used CIFAR-10 is, in fact, a limited data benchmark, and improve the record FID from 5.
  • 9 Jul 2018 27 repositories listed Syntology ran 75 of 129 samples · 54 unverified · 44 pointer-only (licence)
    Flow-based generative models (Dinh et al., 2014) are conceptually attractive due to tractability of the exact log-likelihood, tractability of exact latent-variable inference, and parallelizability of both training and…
  • 18 Mar 2019 24 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 2 pointer-only (licence)
    Previous methods directly feed the semantic layout as input to the deep network, which is then processed through stacks of convolution, normalization, and nonlinearity layers.
  • 12 Feb 2018 22 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)
    Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales.
  • 27 Jul 2016 22 repositories listed Syntology ran 1 of 18 samples · 17 unverified
    It this paper we revisit the fast stylization method introduced in Ulyanov et.
  • 11 May 2021 21 repositories listed Syntology ran 27 of 50 samples · 23 unverified · 16 pointer-only (licence)
    Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.
  • 30 Nov 2017 21 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 3 pointer-only (licence)
    We present a new method for synthesizing high-resolution photo-realistic images from semantic label maps using conditional generative adversarial networks (conditional GANs).
  • 10 Dec 2016 21 repositories listed Syntology ran 11 of 31 samples · 20 unverified · 1 pointer-only (licence)
    Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications.
  • 1 Jun 2022 20 repositories listed Syntology ran 11 of 21 samples · 10 unverified · 13 pointer-only (licence)
    We argue that the theory and practice of diffusion-based generative models are currently unnecessarily convoluted and seek to remedy the situation by presenting a design space that clearly separates the concrete design…
  • 28 Nov 2017 20 repositories listed Syntology ran 3 of 3 samples · 0 unverified
    In this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation.
  • 25 Jan 2016 20 repositories listed Syntology ran 19 of 29 samples · 10 unverified · 18 pointer-only (licence)
    Modeling the distribution of natural images is a landmark problem in unsupervised learning.
  • 16 Feb 2015 20 repositories listed Syntology ran 7 of 18 samples · 11 unverified · 9 pointer-only (licence)
    This paper introduces the Deep Recurrent Attentive Writer (DRAW) neural network architecture for image generation.

Syntology lines on 30 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections