Datasets › Multi-Modal CelebA-HQ

Multi-Modal CelebA-HQ

Introduced by Weihao Xia et al. in TediGAN: Text-Guided Diverse Face Image Generation and Manipulation6 Dec 2020 archive 2025-07-28

Multi-Modal-CelebA-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ. Each image has high-quality segmentation mask, sketch, descriptive text, and image with transparent background.

Multi-Modal-CelebA-HQ can be used to train and evaluate algorithms of text-to-image-generation, text-guided image manipulation, sketch-to-image generation, and GANs for face generation and editing.

Source: Multi-Modal CelebA-HQ Dataset Image Source: Xia et al

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-to-Image Generation Multi-Modal-CelebA-HQ Swinv2-Imagen FID 10.31 Swinv2-Imagen: Hierarchical Vision Transformer Diffusion... — 10 Compare
Face Sketch Synthesis Multi-Modal CelebA-HQ Diffusion FID 26.09 Unite and Conquer: Plug & Play Multi-Modal Synthesis... Nithin-GK/UniteandConquer 1 Compare
multimodal generation Multi-Modal CelebA-HQ Diffusion FID 26.09 Unite and Conquer: Plug & Play Multi-Modal Synthesis... Nithin-GK/UniteandConquer 1 Compare

Papers archive 2025-07-28

10 shown of 10 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 27. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Unite and Conquer: Plug & Play Multi-Modal Synthesis using Diffusion Models 1 3 1 Dec 2022 not harvested
Shifted Diffusion for Text-to-image Generation 1 1 24 Nov 2022 ran 7 of 17 samples (10 unverified)
Swinv2-Imagen: Hierarchical Vision Transformer Diffusion Models for Text-to-Image Generation 0 1 18 Oct 2022 not harvested
LAFITE: Towards Language-Free Training for Text-to-Image Generation 3 1 27 Nov 2021 ran 4 of 18 samples (14 unverified)
Towards Open-World Text-Guided Face Image Generation and Manipulation 2 1 18 Apr 2021 ran 2 of 5 samples (3 unverified)
TediGAN: Text-Guided Diverse Face Image Generation and Manipulation 5 1 6 Dec 2020 not harvested
DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis 3 1 13 Aug 2020 not harvested
Controllable Text-to-Image Generation 2 1 16 Sep 2019 ran 0 of 10 samples (10 unverified)
DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis 4 1 2 Apr 2019 ran 4 of 5 samples (1 unverified; 1 pointer-only for licence)
AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks 20 1 28 Nov 2017 ran 3 of 3 samples (0 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Multi-Modal-CelebA-HQ
  • Multi-Modal CelebA-HQ

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections