Home › Datasets › task › Image Classification

Image Classification datasets

archive 2025-07-28

283 datasets carry the task tag "Image Classification" (the task itself: Image Classification), ordered by the archive's paper count. Page 1 of 6: 48 shown of 283. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Image Classification datasets 1–48 of 283

description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
CelebA (CelebFaces Attributes Dataset)
CelebFaces Attributes dataset contains 202,599 face images of the size 178×218 from 10,177 celebrities, each annotated with 40 binary labels indicating facial attributes like hair color, gender and age.
3,477 papers · 17 benchmarks
SVHN (Street View House Numbers)
Street View House Numbers (SVHN) is a digit classification benchmark dataset that contains 600,000 32×32 RGB images of printed digits (from 0 to 9) cropped from pictures of house number plates.
3,406 papers · 12 benchmarks
Fashion-MNIST is a dataset comprising of 28×28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category.
3,202 papers · 15 benchmarks
CUB-200-2011 (Caltech-UCSD Birds-200-2011)
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
Oxford 102 Flower (102 Category Flower Dataset)
Oxford 102 Flower is an image classification dataset consisting of 102 flower categories.
1,307 papers · 16 benchmarks
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images.
1,232 papers · 7 benchmarks
STL-10 (Self-Taught Learning 10)
The STL-10 is an image dataset derived from ImageNet and popularly used to evaluate algorithms of unsupervised feature learning or self-taught learning.
1,092 papers · 18 benchmarks
DTD (Describable Textures Dataset)
The Describable Textures Dataset (DTD) contains 5640 texture images in the wild.
870 papers · 8 benchmarks
The Food-101 dataset consists of 101 food categories with 750 training and 250 test images per category, making a total of 101k images.
805 papers · 14 benchmarks
The Stanford Cars dataset consists of 196 classes of cars with a total of 16,185 images, taken from the rear.
790 papers · 13 benchmarks
Eurosat is a dataset and deep learning benchmark for land use and land cover classification.
687 papers · 8 benchmarks
The iNaturalist 2017 dataset (iNat) contains 675,170 training and validation images from 5,089 natural fine-grained categories.
603 papers · 12 benchmarks
The Places205 dataset is a large-scale scene-centric dataset with 205 common scene categories.
525 papers · 1 benchmark
FGVC-Aircraft contains 10,200 images of aircraft, with 100 images for each of 102 different aircraft model variants, most of which are airplanes.
520 papers · 12 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
Caltech-256 is an object recognition dataset containing 30,607 real-world images, of different sizes, spanning 257 classes (256 object classes and an additional clutter class).
401 papers · 4 benchmarks
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
387 papers · 4 benchmarks
GTSRB (German Traffic Sign Recognition Benchmark)
The German Traffic Sign Recognition Benchmark (GTSRB) contains 43 classes of traffic signs, split into 39,209 training images and 12,630 test images.
374 papers · 5 benchmarks
The tieredImageNet dataset is a larger subset of ILSVRC-12 with 608 classes (779,165 images) grouped into 34 higher-level nodes in the ImageNet human-curated hierarchy.
317 papers · 7 benchmarks
Clothing1M contains 1M clothing images in 14 classes.
288 papers · 4 benchmarks
ImageNet-Sketch data set consists of 50,889 images, approximately 50 images for each of the 1000 ImageNet classes.
268 papers · 3 benchmarks
EMNIST (Extended MNIST)
EMNIST (extended MNIST) has 4 times more data than MNIST.
264 papers · 10 benchmarks
NAS-Bench-201 is a benchmark (and search space) for neural architecture search.
260 papers · 4 benchmarks
YFCC100M is a that dataset contains a total of 100 million media objects, of which approximately 99.2 million are photos and 0.8 million are videos, all of which carry a Creative Commons license.
243 papers · 0 benchmarks
Stanford Online Products (SOP) dataset has 22,634 classes with 120,053 product images.
231 papers · 5 benchmarks
AI2D (AI2 Diagrams)
AI2 Diagrams (AI2D) is a dataset of over 5000 grade school science diagrams with over 150000 rich annotations, their ground truth syntactic parses, and more than 15000 corresponding multiple choice questions.
207 papers · 1 benchmark
VTAB (Visual Task Adaptation Benchmark)
The Visual Task Adaptation Benchmark (VTAB) is a benchmark designed to evaluate general visual representations².
202 papers · 4 benchmarks
CINIC-10 is a dataset for image classification.
197 papers · 3 benchmarks
RESISC45 dataset is a dataset for Remote Sensing Image Scene Classification (RESISC).
187 papers · 3 benchmarks
The Extended Yale B database contains 2414 frontal-face images with size 192×168 over 38 subjects and about 64 images per subject.
185 papers · 1 benchmark
The WebVision dataset is designed to facilitate the research on learning visual representation from noisy web data.
179 papers · 4 benchmarks
LabelMe database is a large collection of images with ground truth labels for object detection and recognition.
178 papers · 1 benchmark
ObjectNet is a test set of images collected directly using crowd-sourcing.
155 papers · 4 benchmarks
fMoW (Functional Map of the World)
Functional Map of the World (fMoW) is a dataset that aims to inspire the development of machine learning models capable of predicting the functional purpose of buildings and land use from temporal sequences of satellite images and a rich…
144 papers · 1 benchmark
Oxford5k (Oxford Buildings)
Oxford5K is the Oxford Buildings Dataset, which contains 5062 images collected from Flickr.
137 papers · 1 benchmark
The Meta-Dataset benchmark is a large few-shot learning benchmark and consists of multiple datasets of different data distributions.
128 papers · 2 benchmarks
JFT-300M is an internal Google dataset used for training image classification models.
123 papers · 1 benchmark
Kvasir (The Kvasir Dataset)
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
Imagenet32 is a huge dataset made up of small images called the down-sampled version of Imagenet.
112 papers · 4 benchmarks
The smallNORB dataset is a datset for 3D object recognition from shape.
112 papers · 1 benchmark
N-Caltech 101 (Neuromorphic-Caltech101)
The Neuromorphic-Caltech101 (N-Caltech101) dataset is a spiking version of the original frame-based Caltech101 dataset.
110 papers · 3 benchmarks
PCam (PatchCamelyon)
PatchCamelyon is an image classification dataset.
110 papers · 4 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.