Home › Datasets › task › Classification
Classification datasets
archive 2025-07-28
222 datasets carry the task tag "Classification" (the task itself: Classification), ordered by the archive's paper count. Page 1 of 5: 48 shown of 222. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Classification datasets 1–48 of 222
description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
GLUE (General Language Understanding Evaluation benchmark)
General Language Understanding Evaluation (GLUE) benchmark is a collection of nine natural language understanding tasks, including single-sentence tasks CoLA and SST-2, similarity and paraphrasing tasks MRPC, STS-B and QQP, and natural…
3,197 papers · 13 benchmarks
SST (Stanford Sentiment Treebank)
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language.
2,354 papers · 6 benchmarks
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language.
1,808 papers · 2 benchmarks
DTD (Describable Textures Dataset)
The Describable Textures Dataset (DTD) contains 5640 texture images in the wild.
870 papers · 8 benchmarks
The Food-101 dataset consists of 101 food categories with 750 training and 250 test images per category, making a total of 101k images.
805 papers · 14 benchmarks
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples.
701 papers · 5 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
Eurosat is a dataset and deep learning benchmark for land use and land cover classification.
687 papers · 8 benchmarks
Common corruptions dataset for CIFAR10
494 papers · 2 benchmarks
WSC (Winograd Schema Challenge)
The Winograd Schema Challenge was introduced both as an alternative to the Turing Test and as a test of a system’s ability to do commonsense reasoning.
361 papers · 2 benchmarks
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
SGD (Schema-Guided Dialogue)
The Schema-Guided Dialogue (SGD) dataset consists of over 20k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
186 papers · 2 benchmarks
fMoW (Functional Map of the World)
Functional Map of the World (fMoW) is a dataset that aims to inspire the development of machine learning models capable of predicting the functional purpose of buildings and land use from temporal sequences of satellite images and a rich…
144 papers · 1 benchmark
The Neuromorphic-Caltech101 (N-Caltech101) dataset is a spiking version of the original frame-based Caltech101 dataset.
110 papers · 3 benchmarks
The data is related with direct marketing campaigns (phone calls) of a Portuguese banking institution.
69 papers · 0 benchmarks
CIFAR-10H is a new dataset of soft labels reflecting human perceptual uncertainty for the 10,000-image CIFAR-10 test set.
61 papers · 0 benchmarks
Data Set Information: Extraction was done by Barry Becker from the 1994 Census database.
56 papers · 2 benchmarks
A large real-world event-based dataset for object classification.
56 papers · 2 benchmarks
RTE (Recognizing Textual Entailment)
The Recognizing Textual Entailment (RTE) datasets come from a series of textual entailment challenges.
56 papers · 2 benchmarks
HRF (High-Resolution Fundus)
The HRF dataset is a dataset for retinal vessel segmentation which comprises 45 images and is organized as 15 subsets.
53 papers · 3 benchmarks
Contains hundreds of frontal view X-rays and is the largest public resource for COVID-19 image and prognostic data, making it a necessary resource to develop and evaluate tools to aid in the treatment of COVID-19.
35 papers · 1 benchmark
MHIST (Minimalist Histopathology image analysis dataset)
The minimalist histopathology image analysis dataset (MHIST) is a binary classification dataset of 3,152 fixed-size images of colorectal polyps, each with a gold-standard label determined by the majority vote of seven board-certified…
28 papers · 1 benchmark
Pulsar candidates collected during the HTRU survey.
27 papers · 0 benchmarks
TCGA (The Cancer Genome Atlas)
23 papers · 2 benchmarks
A team of researchers from Qatar University, Doha, Qatar, and the University of Dhaka, Bangladesh along with their collaborators from Pakistan and Malaysia in collaboration with medical doctors have created a database of chest X-ray images…
21 papers · 0 benchmarks
he RSSCN7 dataset contains satellite images acquired from Google Earth, which is originally collected for remote sensing scene classification.
20 papers · 1 benchmark
The purpose of this dataset was to study gender bias in occupations.
19 papers · 1 benchmark
Data was collected for normal bearings, single-point drive end and fan end defects.
15 papers · 1 benchmark
N-ImageNet (Large-Scale Dataset for Event-Based Object Recognition)
The N-ImageNet dataset is an event-camera counterpart for the ImageNet dataset.
15 papers · 2 benchmarks
The goal for ISIC 2019 is classify dermoscopic images among nine different diagnostic categories.25,331 images are available for training across 8 different categories.
13 papers · 3 benchmarks
The SD-198 dataset contains 198 different diseases from different types of eczema, acne and various cancerous conditions.
11 papers · 0 benchmarks
GRAZPEDWRI-DX is a public dataset of 20,327 pediatric wrist trauma X-ray images released by the University of Medicine of Graz.
10 papers · 4 benchmarks
Dataset Introduction In this work, we introduce the In-Diagram Logic (InDL) dataset, an innovative resource crafted to rigorously evaluate the logic interpretation abilities of deep learning models.
10 papers · 1 benchmark
SST-3 (Stanford Sentiment Treebank: 3-way)
SST-5 is the Stanford Sentiment Treebank 5-way classification dataset (positive, somewhat positive, neutral, somewhat negative, negative).
10 papers · 1 benchmark
SciRepEval is a comprehensive benchmark for training and evaluating scientific document representations.
9 papers · 0 benchmarks
ArtiFact (Artificial and Factual Image Dataset for Synthetic Image Detection)
The ArtiFact dataset is a large-scale image dataset that aims to include a diverse collection of real and synthetic images from multiple categories, including Human/Human Faces, Animal/Animal Faces, Places, Vehicles, Art, and many other…
7 papers · 0 benchmarks
This dataset is a combination of the following three datasets : figshare, SARTAJ dataset and Br35H This dataset contains 7022 images of human brain MRI images which are classified into 4 classes: glioma - meningioma - no tumor and…
7 papers · 3 benchmarks
SSC (Spiking Speech Commands v0.2)
The SSC dataset is a spiking version of the Speech Commands dataset release by Google (Speech Commands).
7 papers · 1 benchmark
A public data set of walking full-body kinematics and kinetics in individuals with Parkinson’s disease
6 papers · 1 benchmark
RxRx1 is a biological dataset designed specifically for the systematic study of batch effect correction methods.
6 papers · 1 benchmark
XImageNet-12 (XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation)
Enlarge the dataset to understand how image background effect the Computer Vision ML model.
6 papers · 1 benchmark
We construct the ForgeryNet dataset, an extremely large face forgery dataset with unified annotations in image- and video-level data across four tasks: 1) Image Forgery Classification, including two-way (real / fake), three-way (real /…
5 papers · 1 benchmark
RITE (Retinal Images vessel Tree Extraction)
The RITE (Retinal Images vessel Tree Extraction) is a database that enables comparative studies on segmentation or classification of arteries and veins on retinal fundus images, which is established based on the public available DRIVE…
5 papers · 2 benchmarks
VNAT (VPN/NONVPN NETWORK APPLICATION TRAFFIC DATASET)
This dataset is a collection of labelled PCAP files, both encrypted and unencrypted, across 10 applications, as well as a pandas dataframe in HDF5 format containing detailed metadata summarizing the connections from those files.
5 papers · 0 benchmarks
This dataset is described in the ALTA 2021 Shared Task website and associated CodaLab competition.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.