Home › Datasets › task › Image Clustering
Image Clustering datasets
archive 2025-07-28
44 datasets carry the task tag "Image Clustering" (the task itself: Image Clustering), ordered by the archive's paper count. Page 1 of 1: 44 shown of 44. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Image Clustering datasets 1–44 of 44
description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
Fashion-MNIST is a dataset comprising of 28×28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category.
3,202 papers · 15 benchmarks
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
Oxford 102 Flower is an image classification dataset consisting of 102 flower categories.
1,307 papers · 16 benchmarks
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images.
1,232 papers · 7 benchmarks
STL-10 (Self-Taught Learning 10)
The STL-10 is an image dataset derived from ImageNet and popularly used to evaluate algorithms of unsupervised feature learning or self-taught learning.
1,092 papers · 18 benchmarks
DTD (Describable Textures Dataset)
The Describable Textures Dataset (DTD) contains 5640 texture images in the wild.
870 papers · 8 benchmarks
The Food-101 dataset consists of 101 food categories with 750 training and 250 test images per category, making a total of 101k images.
805 papers · 14 benchmarks
The Stanford Cars dataset consists of 196 classes of cars with a total of 16,185 images, taken from the rear.
790 papers · 13 benchmarks
The Caltech101 dataset contains images from 101 object categories (e.g., “helicopter”, “elephant” and “chair” etc.) and a background category that contains the images not from the 101 object categories.
709 papers · 10 benchmarks
Eurosat is a dataset and deep learning benchmark for land use and land cover classification.
687 papers · 8 benchmarks
FGVC-Aircraft contains 10,200 images of aircraft, with 100 images for each of 102 different aircraft model variants, most of which are airplanes.
520 papers · 12 benchmarks
USPS is a digit dataset automatically scanned from envelopes by the U.S.
459 papers · 2 benchmarks
GTSRB (German Traffic Sign Recognition Benchmark)
The German Traffic Sign Recognition Benchmark (GTSRB) contains 43 classes of traffic signs, split into 39,209 training images and 12,630 test images.
374 papers · 5 benchmarks
HAR (Human Activity Recognition Using Smartphones)
The Human Activity Recognition Dataset has been collected from 30 subjects performing six different activities (Walking, Walking Upstairs, Walking Downstairs, Sitting, Standing, Laying).
307 papers · 3 benchmarks
EMNIST (extended MNIST) has 4 times more data than MNIST.
264 papers · 10 benchmarks
RESISC45 dataset is a dataset for Remote Sensing Image Scene Classification (RESISC).
187 papers · 3 benchmarks
The Extended Yale B database contains 2414 frontal-face images with size 192×168 over 38 subjects and about 64 images per subject.
185 papers · 1 benchmark
The Hateful Memes data set is a multimodal dataset for hateful meme detection (image + text) that contains 10,000+ new multimodal examples created by Facebook AI.
177 papers · 3 benchmarks
FER2013 (Facial Expression Recognition 2013 Dataset)
Fer2013 contains approximately 30,000 facial RGB images of different expressions with size restricted to 48×48, and the main labels of it can be divided into 7 types: 0=Angry, 1=Disgust, 2=Fear, 3=Happy, 4=Sad, 5=Surprise, 6=Neutral.
168 papers · 5 benchmarks
PatchCamelyon is an image classification dataset.
110 papers · 4 benchmarks
FRGC (Face Recognition Grand Challenge)
The data for FRGC consists of 50,000 recordings divided into training and validation partitions.
102 papers · 1 benchmark
Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes.
95 papers · 3 benchmarks
Birdsnap is a large bird dataset consisting of 49,829 images from 500 bird species with 47,386 images used for training and 2,443 images used for testing.
72 papers · 2 benchmarks
The Oxford-IIIT Pet Dataset is a 37-category pet dataset with roughly 200 images for each class.
59 papers · 5 benchmarks
The Stanford Dogs dataset contains 20,580 images of 120 classes of dogs from around the world, which are divided into 12,000 images for training and 8,580 images for testing.
57 papers · 6 benchmarks
The Scene UNderstanding (SUN) database contains 899 categories and 130,519 images.
52 papers · 8 benchmarks
Letter (Letter Recognition Data Set)
Letter Recognition Data Set is a handwritten digit dataset.
49 papers · 2 benchmarks
CARS196 is composed of 16,185 car images of 196 classes.
43 papers · 4 benchmarks
This split was introduced in TEMI (BMVC 2023) Adaloglou, Nikolas, Felix Michels, Hamza Kalisch, and Markus Kollmann.
12 papers · 4 benchmarks
Fashion 144K is a novel heterogeneous dataset with 144,169 user posts containing diverse image, textual and meta information.
11 papers · 0 benchmarks
The ImageNet-50 dataset split as introduced in TEMI.
4 papers · 1 benchmark
Country211 is a dataset released by OpenAI, designed to assess the geolocation capability of visual representations.
3 papers · 2 benchmarks
The Sheffield (previously UMIST) Face Database consists of 564 images of 20 individuals (mixed race/gender/appearance).
3 papers · 1 benchmark
The Rendered SST2 dataset is a dataset released by OpenAI, that measures the optical character recognition capability of visual representations.
2 papers · 1 benchmark
Card is a dataset of playing card images, which consists of 8,029 images with two clusterings, i.e., rank (Ace, King, Queen, etc.) and suits (clubs, diamonds, hearts, spades).
1 paper · 0 benchmarks
The data consists of 21 images of microtubules in PFA-fixed NIH 3T3 mouse embryonic fibroblasts (DSMZ: ACC59) labeled with a mouse anti-alpha-tubulin monoclonal IgG1 antibody (Thermofisher A11126, primary antibody) and visualized by a…
1 paper · 0 benchmarks
SPOT-10 (Animal Pattern Benchmark Dataset for Machine Learning Algorithms)
The SPOTS-10 dataset is an extensive collection of grayscale images showcasing diverse patterns found in ten animal species.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.