Home › Datasets › task › Image Classification
Image Classification datasets
archive 2025-07-28
283 datasets carry the task tag "Image Classification" (the task itself: Image Classification), ordered by the archive's paper count. Page 2 of 6: 48 shown of 283. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Image Classification datasets 49–96 of 283
The Stylized-ImageNet dataset is created by removing local texture cues in ImageNet while retaining global shape information on natural images via AdaIN style transfer.
106 papers · 1 benchmark
Comprises 11 hand gesture categories from 29 subjects under 3 illumination conditions.
103 papers · 6 benchmarks
The image dataset TinyImages contains 80 million images of size 32×32 collected from the Internet, crawling the words in WordNet.
103 papers · 0 benchmarks
This work presents two new benchmark datasets (CIFAR-10N, CIFAR-100N), equipping the training dataset of CIFAR-10 and CIFAR-100 with human-annotated real-world noisy labels that we collect from Amazon Mechanical Turk.
97 papers · 6 benchmarks
Kuzushiji-MNIST is a drop-in replacement for the MNIST dataset (28x28 grayscale, 70,000 images).
97 papers · 2 benchmarks
The Oxford-IIIT Pet Dataset has 37 categories with roughly 200 images for each class.
90 papers · 5 benchmarks
ImageNet-O consists of images from classes that are not found in the ImageNet-1k dataset.
89 papers · 0 benchmarks
AbstractReasoning is a dataset for abstract reasoning, where the goal is to infer the correct answer from the context panels based on abstract reasoning.
86 papers · 0 benchmarks
BigEarthNet consists of 590,326 Sentinel-2 image patches, each of which is a section of i) 120x120 pixels for 10m bands; ii) 60x60 pixels for 20m bands; and iii) 20x20 pixels for 60m bands.
85 papers · 3 benchmarks
ChestX-ray8 is a medical imaging dataset which comprises 108,948 frontal-view X-ray images of 32,717 (collected from the year of 1992 to 2015) unique patients with the text-mined eight common disease labels, mined from the text…
81 papers · 0 benchmarks
The PlantVillage dataset consists of 54303 healthy and unhealthy leaf images divided into 38 categories by species and disease.
68 papers · 1 benchmark
This work presents two new benchmark datasets (CIFAR-10N, CIFAR-100N), equipping the training dataset of CIFAR-10 and CIFAR-100 with human-annotated real-world noisy labels that we collect from Amazon Mechanical Turk.
66 papers · 1 benchmark
The Places365 dataset is a scene recognition dataset.
65 papers · 7 benchmarks
The Oxford-IIIT Pet Dataset is a 37-category pet dataset with roughly 200 images for each class.
59 papers · 5 benchmarks
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
PGM (Procedurally Generated Matrices (PGM))
PGM dataset serves as a tool for studying both abstract reasoning and generalisation in models.
57 papers · 0 benchmarks
MINC (Materials in Context Database)
MINC is a large-scale, open dataset of materials in the wild.
56 papers · 0 benchmarks
Tiny ImageNet-C is an open-source data set comprising algorithmically generated corruptions applied to the Tiny ImageNet (ImageNet-200) test set comprising 200 classes following the concept of ImageNet-C.
55 papers · 0 benchmarks
The MultiMNIST dataset is generated from MNIST.
52 papers · 1 benchmark
The Scene UNderstanding (SUN) database contains 899 categories and 130,519 images.
52 papers · 8 benchmarks
CARS196 is composed of 16,185 car images of 196 classes.
43 papers · 4 benchmarks
JFT-3B is an internal Google dataset and a larger version of the JFT-300M dataset.
41 papers · 0 benchmarks
Million-AID is a large-scale benchmark dataset containing a million instances for RS scene classification.
41 papers · 0 benchmarks
Visual Wake Words represents a common microcontroller vision use-case of identifying whether a person is present in the image or not, and provides a realistic benchmark for tiny vision models.
40 papers · 1 benchmark
Imagenette is a subset of 10 easily classified classes from Imagenet (bench, English springer, cassette player, chain saw, church, French horn, garbage truck, gas pump, golf ball, parachute).
38 papers · 1 benchmark
The SUN Attribute dataset consists of 14,340 images from 717 scene categories, and each category is annotated with a taxonomy of 102 discriminate attributes.
38 papers · 2 benchmarks
Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes.
37 papers · 1 benchmark
Imagenet64 is a massive dataset of small images called the down-sampled version of Imagenet.
35 papers · 1 benchmark
ImageNet-P consists of noise, blur, weather, and digital distortions.
32 papers · 1 benchmark
Omni-Realm Benchmark (OmniBenchmark) is a diverse (21 semantic realm-wise datasets) and concise (realm-wise datasets have no concepts overlapping) benchmark for evaluating pre-trained model generalization across semantic…
29 papers · 1 benchmark
Dataset with 28,792 retinal images from the EyePACS dataset, based on a three-level quality grading system (i.e., Good', Usable' and Reject') for evaluating RIQA methods.
28 papers · 0 benchmarks
iNat2021 is a large-scale image dataset collected and annotated by community scientists that contains over 2.7M images from 10k different species.
27 papers · 0 benchmarks
Object detection benchmark for logo detection.
26 papers · 3 benchmarks
The exact pre-processing steps used to construct the MNIST dataset have long been lost.
26 papers · 2 benchmarks
UrbanCars facilitates multi-shortcut learning under the controlled setting with two shortcuts—background and co-occurring object.
26 papers · 0 benchmarks
BAM! (Behance Artistic Media)
The Behance Artistic Media dataset (BAM!) is a large-scale dataset of contemporary artwork from Behance, a website containing millions of portfolios from professional and commercial artists.
25 papers · 0 benchmarks
ELEVATER (Evaluation of Language-augmented Visual Task-level Transfer)
The ELEVATER benchmark is a collection of resources for training, evaluating, and analyzing language-image models on image classification and object detection.
25 papers · 2 benchmarks
ImageNet-W(atermark) is a test set to evaluate models’ reliance on the newly found watermark shortcut in ImageNet, which is used to predict the carton class.
23 papers · 0 benchmarks
IP102 contains more than 75,000 images belonging to 102 categories, which exhibit a natural long-tailed distribution.
22 papers · 0 benchmarks
Our goal is to improve upon the status quo for designing image classification models trained in one domain that perform well on images from another domain.
22 papers · 3 benchmarks
LAG (Large-scale Attention based Glaucoma)
Includes 5,824 fundus images labeled with either positive glaucoma (2,392) or negative glaucoma (3,432).
21 papers · 1 benchmark
A composite dataset that unifies semantic segmentation datasets from different domains.
21 papers · 0 benchmarks
ETHOS (multi-labEl haTe speecH detectiOn dataSet)
ETHOS is a hate speech detection dataset.
20 papers · 2 benchmarks
PlantDoc is a dataset for visual plant disease detection.
20 papers · 1 benchmark
Chaoyang dataset contains 1111 normal, 842 serrated, 1404 adenocarcinoma, 664 adenoma, and 705 normal, 321 serrated, 840 adenocarcinoma, 273 adenoma samples for training and testing, respectively.
19 papers · 2 benchmarks
DeepFish as a benchmark suite with a large-scale dataset to train and test methods for several computer vision tasks.
19 papers · 1 benchmark
An annotated image memorability dataset to date (with 60,000 labeled images from a diverse array of sources).
18 papers · 0 benchmarks
MLRSNet is a a multi-label high spatial resolution remote sensing dataset for semantic scene understanding.
17 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.