Home › Datasets › task › Image Classification
Image Classification datasets
archive 2025-07-28
283 datasets carry the task tag "Image Classification" (the task itself: Image Classification), ordered by the archive's paper count. Page 3 of 6: 48 shown of 283. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Image Classification datasets 97–144 of 283
See paper: Caldas, Sebastian, et al.
16 papers · 2 benchmarks
Food2K is a large food recognition dataset with 2,000 categories and over 1 million images.
14 papers · 0 benchmarks
Brief Description The Neuromorphic-MNIST (N-MNIST) dataset is a spiking version of the original frame-based MNIST dataset.
14 papers · 4 benchmarks
TAU Urban Acoustic Scenes 2019 development dataset consists of 10-seconds audio segments from 10 acoustic scenes: airport, indoor shopping mall, metro station, pedestrian street, public square, street with medium level of traffic,…
14 papers · 2 benchmarks
A large set of images of cats and dogs.
13 papers · 5 benchmarks
Update on 3DIdent, where we introduce six additional object classes (Hare, Dragon, Cow, Armadillo, Horse, and Head), and impose a causal graph over the latent variables.
12 papers · 1 benchmark
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
This split was introduced in TEMI (BMVC 2023) Adaloglou, Nikolas, Felix Michels, Hamza Kalisch, and Markus Kollmann.
12 papers · 4 benchmarks
BCNB (Early Breast Cancer Core-Needle Biopsy WSI)
Breast cancer (BC) has become the greatest threat to women’s health worldwide.
11 papers · 0 benchmarks
BCN20000 is a dataset composed of 19,424 dermoscopic images of skin lesions captured from 2010 to 2016 in the facilities of the Hospital Clínic in Barcelona.
11 papers · 0 benchmarks
BreakHis (Breast Cancer Histopathological Database)
The Breast Cancer Histopathological Image Classification (BreakHis) is composed of 9,109 microscopic images of breast tumor tissue collected from 82 patients using different magnifying factors (40X, 100X, 200X, and 400X).
11 papers · 5 benchmarks
FoodX-251 is a dataset of 251 fine-grained classes with 118k training, 12k validation and 28k test images.
11 papers · 1 benchmark
Four pathologists from Longhua Hospital Shanghai University of Traditional Chinese Medicine provide 600 images of gastric cancer pathology images at size 2048×2048 pixels.
11 papers · 1 benchmark
So2Sat LCZ42 consists of local climate zone (LCZ) labels of about half a million Sentinel-1 and Sentinel-2 image patches in 42 urban agglomerations (plus 10 additional smaller areas) across the globe.
11 papers · 1 benchmark
Kuzushiji-49 is an MNIST-like dataset that has 49 classes (28x28 grayscale, 270,912 images) from 48 Hiragana characters and one Hiragana iteration mark.
10 papers · 0 benchmarks
AmsterTime (AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain Shift)
AmsterTime dataset offers a collection of 2,500 well-curated images matching the same scene from a street view matched to historical archival image data from Amsterdam city.
9 papers · 3 benchmarks
The iCartoonFace dataset is a large-scale dataset that can be used for two different tasks: cartoon face detection and cartoon face recognition.
9 papers · 1 benchmark
The iWildCam2020-WILDS dataset is a variant of the iWildCam 2020 dataset.
9 papers · 1 benchmark
Dataset aimed to do automated aerial scene classification of disaster events from on-board a UAV.
8 papers · 1 benchmark
ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy.
8 papers · 0 benchmarks
ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy.
8 papers · 2 benchmarks
ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy.
8 papers · 2 benchmarks
The NCT-CRC-HE-100K dataset is a set of 100,000 non-overlapping image patches extracted from 86 H&E stained human cancer tissue slides and normal tissue from the NCT biobank (National Center for Tumor Diseases) and the UMM pathology…
8 papers · 2 benchmarks
We introduce ArtBench-10, the first class-balanced, high-quality, cleanly annotated, and standardized dataset for benchmarking artwork generation.
7 papers · 1 benchmark
Bamboo Dataset is a mega-scale and information-dense dataset for both classification and detection pre-training.
7 papers · 0 benchmarks
A new challenging dataset that can be used for many pattern recognition tasks.
7 papers · 1 benchmark
Danish Fungi 2020 (DF20) is a fine-grained dataset and benchmark.
7 papers · 1 benchmark
The Diabetic Foot Ulcers dataset (DFUC2021) is a dataset for analysis of pathology, focusing on infection and ischaemia.
7 papers · 0 benchmarks
Grocery Store is a dataset of natural images of grocery items.
7 papers · 0 benchmarks
ImageNet-9 consists of images with different amounts of background and foreground signal, which you can use to measure the extent to which your models rely on image backgrounds.
7 papers · 1 benchmark
The Kannada-MNIST dataset is a drop-in substitute for the standard MNIST dataset for the Kannada language.
7 papers · 0 benchmarks
Consists of faces extracted from pre-modern Japanese artwork.
7 papers · 0 benchmarks
NumtaDB (Assembled Bengali Handwritten Digits)
To benchmark Bengali digit recognition algorithms, a large publicly available dataset is required which is free from biases originating from geographical location, gender, and age.
7 papers · 0 benchmarks
The PMData dataset aims to combine the traditional lifelogging with sports activity logging.
7 papers · 0 benchmarks
The PS-Battles dataset is gathered from a large community of image manipulation enthusiasts and provides a basis for media derivation and manipulation detection in the visual domain.
7 papers · 0 benchmarks
SI-SCORE (Synthetic Interventions on Scenes for Robustness Evaluation)
A synthetic dataset uses for a systematic analysis across common factors of variation.
7 papers · 0 benchmarks
This is a dataset with spurious correlations which can be used to evaluate machine learning methods for out-of-distribution generalization, causal inference, and related field.
6 papers · 1 benchmark
F-CelebA - This dataset is adapted from federated learning.
6 papers · 1 benchmark
The Food-101N dataset is introduced in "CleanNet: Transfer Learning for Scalable Image Training with Label Noise (CVPR'18).
6 papers · 1 benchmark
The dataset contains a total of 27,558 cell images with equal instances of parasitized and uninfected cells.
6 papers · 2 benchmarks
OOD-CV (Out Of Distribution Generalization in Computer Vision)
Enhancing the robustness of vision algorithms in real-world scenarios is challenging.
6 papers · 1 benchmark
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
Part of the Controlled Noisy Web Labels Dataset.
6 papers · 2 benchmarks
Part of the Controlled Noisy Web Labels Dataset.
6 papers · 2 benchmarks
Part of the Controlled Noisy Web Labels Dataset.
6 papers · 2 benchmarks
The Urban Environments dataset is a dataset of 20 land use classes across 300 European cities paired with satellite imagery data.
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.