Home › Datasets › modality › Images

Images datasets

archive 2025-07-28

3,239 datasets carry the modality tag "Images", ordered by the archive's paper count. Page 1 of 68: 48 shown of 3,239. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Images datasets 1–48 of 3,239

description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
CelebA (CelebFaces Attributes Dataset)
CelebFaces Attributes dataset contains 202,599 face images of the size 178×218 from 10,177 celebrities, each annotated with 40 binary labels indicating facial attributes like hair color, gender and age.
3,477 papers · 17 benchmarks
SVHN (Street View House Numbers)
Street View House Numbers (SVHN) is a digit classification benchmark dataset that contains 600,000 32×32 RGB images of printed digits (from 0 to 9) cropped from pictures of house number plates.
3,406 papers · 12 benchmarks
Fashion-MNIST is a dataset comprising of 28×28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category.
3,202 papers · 15 benchmarks
CUB-200-2011 (Caltech-UCSD Birds-200-2011)
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
ShapeNet is a large scale repository for 3D CAD models developed by researchers from Stanford University, Princeton University and the Toyota Technological Institute at Chicago, USA.
1,947 papers · 13 benchmarks
Visual Question Answering (VQA) is a dataset containing open-ended questions about images.
1,834 papers · 0 benchmarks
ScanNet is an instance-level indoor RGB-D dataset that includes both 2D and 3D data.
1,595 papers · 21 benchmarks
FFHQ (Flickr-Faces-HQ)
Flickr-Faces-HQ (FFHQ) consists of 70,000 high-quality PNG images at 1024×1024 resolution and contains considerable variation in terms of age, ethnicity and image background.
1,468 papers · 17 benchmarks
Oxford 102 Flower (102 Category Flower Dataset)
Oxford 102 Flower is an image classification dataset consisting of 102 flower categories.
1,307 papers · 16 benchmarks
Visual Genome contains Visual Question Answering data in a multi-choice setting.
1,256 papers · 15 benchmarks
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images.
1,232 papers · 7 benchmarks
The Places dataset is proposed for scene recognition and contains more than 2.5 million images covering more than 205 scene categories with more than 5,000 images per category.
1,151 papers · 4 benchmarks
STL-10 (Self-Taught Learning 10)
The STL-10 is an image dataset derived from ImageNet and popularly used to evaluate algorithms of unsupervised feature learning or self-taught learning.
1,092 papers · 18 benchmarks
Office-Home is a benchmark dataset for domain adaptation which contains 4 domains where each domain consists of 65 categories.
1,074 papers · 11 benchmarks
The CelebA-HQ dataset is a high-quality version of CelebA that consists of 30,000 images at 1024×1024 resolution.
954 papers · 13 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
Market-1501 is a large-scale public benchmark dataset for person re-identification.
873 papers · 9 benchmarks
DTD (Describable Textures Dataset)
The Describable Textures Dataset (DTD) contains 5640 texture images in the wild.
870 papers · 8 benchmarks
LSUN (Large-scale Scene UNderstanding Challenge)
The Large-scale Scene Understanding (LSUN) challenge aims to provide a different benchmark for large-scale scene classification and understanding.
867 papers · 10 benchmarks
LFW (Labeled Faces in the Wild)
The LFW dataset contains 13,233 images of faces collected from the web.
820 papers · 12 benchmarks
The Food-101 dataset consists of 101 food categories with 750 training and 250 test images per category, making a total of 101k images.
805 papers · 14 benchmarks
The Stanford Cars dataset consists of 196 classes of cars with a total of 16,185 images, taken from the rear.
790 papers · 13 benchmarks
The Human3.6M dataset is one of the largest motion capture datasets, which consists of 3.6 million human poses and corresponding images captured by a high-speed motion capture system.
783 papers · 13 benchmarks
The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs.
749 papers · 8 benchmarks
BSD (Berkeley Segmentation Dataset)
BSD is a dataset used frequently for image denoising and super-resolution.
718 papers · 48 benchmarks
The dataset contains 400 human action classes, with at least 400 video clips for each action.
712 papers · 0 benchmarks
The Caltech101 dataset contains images from 101 object categories (e.g., “helicopter”, “elephant” and “chair” etc.) and a background category that contains the images not from the 101 object categories.
709 papers · 10 benchmarks
Eurosat is a dataset and deep learning benchmark for land use and land cover classification.
687 papers · 8 benchmarks
PACS (Photo-Art-Cartoon-Sketch)
PACS is an image dataset for domain generalization.
668 papers · 10 benchmarks
CLEVR (Compositional Language and Elementary Visual Reasoning)
CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset.
657 papers · 3 benchmarks
DIV2K is a popular single-image super-resolution dataset which contains 1,000 images with different scenes and is splitted to 800 for training, 100 for validation and 100 for testing.
654 papers · 3 benchmarks
Office-31 (Office Dataset)
The Office dataset contains 31 object categories in three domains: Amazon, DSLR and Webcam.
643 papers · 7 benchmarks
The CheXpert dataset contains 224,316 chest radiographs of 65,240 patients with both frontal and lateral views available.
628 papers · 3 benchmarks
The iNaturalist 2017 dataset (iNat) contains 675,170 training and validation images from 5,089 natural fine-grained categories.
603 papers · 12 benchmarks
ImageNet-C is an open source data set that consists of algorithmically generated corruptions (blur, noise) applied to the ImageNet test-set.
602 papers · 4 benchmarks
VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media.
564 papers · 5 benchmarks
VGGFace2 (Vggface2: A dataset for recognising faces across pose and age)
VGGFace2 is a large-scale face recognition dataset.
539 papers · 3 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.