Home › Datasets › task › Image Generation

Image Generation datasets

archive 2025-07-28

79 datasets carry the task tag "Image Generation" (the task itself: Image Generation), ordered by the archive's paper count. Page 1 of 2: 48 shown of 79. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Image Generation datasets 1–48 of 79

description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
CelebA (CelebFaces Attributes Dataset)
CelebFaces Attributes dataset contains 202,599 face images of the size 178×218 from 10,177 celebrities, each annotated with 40 binary labels indicating facial attributes like hair color, gender and age.
3,477 papers · 17 benchmarks
Fashion-MNIST is a dataset comprising of 28×28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category.
3,202 papers · 15 benchmarks
CUB-200-2011 (Caltech-UCSD Birds-200-2011)
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
FFHQ (Flickr-Faces-HQ)
Flickr-Faces-HQ (FFHQ) consists of 70,000 high-quality PNG images at 1024×1024 resolution and contains considerable variation in terms of age, ethnicity and image background.
1,468 papers · 17 benchmarks
Oxford 102 Flower (102 Category Flower Dataset)
Oxford 102 Flower is an image classification dataset consisting of 102 flower categories.
1,307 papers · 16 benchmarks
STL-10 (Self-Taught Learning 10)
The STL-10 is an image dataset derived from ImageNet and popularly used to evaluate algorithms of unsupervised feature learning or self-taught learning.
1,092 papers · 18 benchmarks
The CelebA-HQ dataset is a high-quality version of CelebA that consists of 30,000 images at 1024×1024 resolution.
954 papers · 13 benchmarks
LSUN (Large-scale Scene UNderstanding Challenge)
The Large-scale Scene Understanding (LSUN) challenge aims to provide a different benchmark for large-scale scene classification and understanding.
867 papers · 10 benchmarks
The Stanford Cars dataset consists of 196 classes of cars with a total of 16,185 images, taken from the rear.
790 papers · 13 benchmarks
CLEVR (Compositional Language and Elementary Visual Reasoning)
CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset.
657 papers · 3 benchmarks
The iNaturalist 2017 dataset (iNat) contains 675,170 training and validation images from 5,089 natural fine-grained categories.
603 papers · 12 benchmarks
Perceptual Similarity is a dataset of human perceptual similarity judgments.
452 papers · 0 benchmarks
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces.
414 papers · 4 benchmarks
FaceForensics++ is a forensics dataset consisting of 1000 original video sequences that have been manipulated with four automated face manipulation methods: Deepfakes, Face2Face, FaceSwap and NeuralTextures.
368 papers · 2 benchmarks
AFHQ (Animal Faces-HQ)
Animal FacesHQ (AFHQ) is a dataset of animal faces consisting of 15,000 high-quality images at 512 × 512 resolution.
327 papers · 6 benchmarks
DensePose (DensePose-COCO)
DensePose-COCO is a large-scale ground-truth dataset with image-to-surface correspondences manually annotated on 50K COCO images and train DensePose-RCNN, to densely regress part-specific UV coordinates within every human region at…
265 papers · 1 benchmark
EMNIST (Extended MNIST)
EMNIST (extended MNIST) has 4 times more data than MNIST.
264 papers · 10 benchmarks
CelebAMask-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
164 papers · 5 benchmarks
ViZDoom is an AI research platform based on the classical First Person Shooter game Doom.
156 papers · 3 benchmarks
LLVIP (A Visible-infrared Paired Dataset for Low-light Vision)
Visible-infrared Paired Dataset for Low-light Vision 30976 images (15488 pairs) 24 dark scenes, 2 daytime scenes Support for image-to-image translation (visible to infrared, or infrared to visible), visible and infrared image fusion,…
116 papers · 6 benchmarks
Structured3D is a large-scale photo-realistic dataset containing 3.5K house designs (a) created by professional designers with a variety of ground truth 3D structure annotations (b) and generate photo-realistic 2D images (c).
85 papers · 7 benchmarks
ARKitScenes is an RGB-D dataset captured with the widely available Apple LiDAR scanner.
75 papers · 2 benchmarks
MetFaces is an image dataset of human faces extracted from works of art.
75 papers · 2 benchmarks
Recipe1M+ is a dataset which contains one million structured cooking recipes with 13M associated images.
68 papers · 3 benchmarks
The Stanford Dogs dataset contains 20,580 images of 120 classes of dogs from around the world, which are divided into 12,000 images for training and 8,580 images for testing.
57 papers · 6 benchmarks
VLN-CE (Vision-and-Language Navigation in Continuous Environments)
Vision and Language Navigation in Continuous Environments (VLN-CE) is an instruction-guided navigation task with crowdsourced instructions, realistic environments, and unconstrained agent navigation.
55 papers · 1 benchmark
The Stacked MNIST dataset is derived from the standard MNIST dataset with an increased number of discrete modes.
43 papers · 1 benchmark
Fashion-Gen consists of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists.
36 papers · 0 benchmarks
GVGAI (General Video Game AI)
The General Video Game AI (GVGAI) framework is widely used in research which features a corpus of over 100 single-player games and 60 two-player games.
36 papers · 0 benchmarks
UT Zappos50K is a large shoe dataset consisting of 50,025 catalog images collected from Zappos.com.
32 papers · 2 benchmarks
LHQ (Landscapes High-Quality)
A dataset of 90,000 high-resolution nature landscape images, crawled from Unsplash and Flickr and preprocessed with Mask R-CNN and Inception V3.
27 papers · 4 benchmarks
Multi-Modal-CelebA-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
27 papers · 3 benchmarks
MMAct is a large-scale dataset for multi/cross modal action understanding.
23 papers · 1 benchmark
A simulation-based dataset featuring 20,000 stack configurations composed of a variety of elementary geometric primitives richly annotated regarding semantics and structural stability.
22 papers · 2 benchmarks
WISE, the first benchmark specifically designed for World Knowledge-Informed Semantic Evaluation.
22 papers · 2 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
ActivityNet-Entities, augments the challenging ActivityNet Captions dataset with 158k bounding box annotations, each grounding a noun phrase.
18 papers · 0 benchmarks
A binarized version of MNIST.
10 papers · 1 benchmark
A large-scale human image dataset with over 230K samples capturing diverse poses and textures.
10 papers · 0 benchmarks
9 papers · 2 benchmarks
We introduce ArtBench-10, the first class-balanced, high-quality, cleanly annotated, and standardized dataset for benchmarking artwork generation.
7 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.