Home › Datasets › task › Classification

Classification datasets

archive 2025-07-28

222 datasets carry the task tag "Classification" (the task itself: Classification), ordered by the archive's paper count. Page 5 of 5: 30 shown of 222. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Classification datasets 193–222 of 222

TLF2K (Table-LastFm2K)
Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset.
1 paper · 1 benchmark
TML1M (Table-MovieLens1M)
Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset.
1 paper · 1 benchmark
Tinto (Tinto: Multisensor Benchmark for 3D Hyperspectral Point Cloud Segmentation in the Geosciences)
The increasing use of deep learning techniques has reduced interpretation time and, ideally, reduced interpreter bias by automatically deriving geological maps from digital outcrop models.
1 paper · 0 benchmarks
Tiny ImageNet-A is a subset of the Tiny ImageNet test set consisting of 3,374 images comprising real-world, unmodified, and naturally occurring examples that are misclassified by ResNet-18.
1 paper · 0 benchmarks
The two Coiling Spiral is a 2d classification dataset composed of two classes; each spiral corresponds to one class.
1 paper · 0 benchmarks
Vulnerable Verified Smart Contracts is a dataset of real vulnerable Ethereum smart contracts.
1 paper · 0 benchmarks
WINGBEATS (MOSQUITO WINGBEAT RECORDINGS)
Context The database contains wav recordings from the same optical sensor inserted in-turn into six insectary boxes containing only one mosquito species of both sexes (about 200-300 flying mosquitoes in each cage).
1 paper · 0 benchmarks
The WORC database consists in total of 930 patients composed of six datasets gathered at the Erasmus MC, consisting of patients with: 1) well-differentiated liposarcoma or lipoma (115 patients); 2) desmoid-type fibromatosis or extremity…
1 paper · 0 benchmarks
The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure.
1 paper · 0 benchmarks
ai4st SLR (Research on AI for Software Testing Research, 2020-2025)
To check the validity of the ai4st ontology, an adapted, lightweight systematic literature review (SLR) was conducted to analyse related research.
1 paper · 0 benchmarks
arXiv Categories (arXiv Categories Multi-label Text Classification Dataset)
This is a dataset of scientific documents derived from arXiv.
1 paper · 0 benchmarks
We release the datasets to replicate the results of Coordinated Reply Attacks in Influence Operations: Characterization and Detection'.
1 paper · 0 benchmarks
uBench (MicroBench)
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks
ALFI (Annotations for Label-Free Images)
ALFI (Annotations for Label-Free Images) is a dataset of images and annotations for label-free microscopy imaging.
0 papers · 0 benchmarks
ALTA 2022 Shared Task (PIBOSO Sentence classification)
This dataset is described in the ALTA 2022 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
ALTA 2023 Shared Task (Discriminate between human-authored and synthetic text generated by Large Language Models (LLMs))
This dataset is described in the ALTA 2023 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware.
0 papers · 0 benchmarks
The dataset consists of 3265 text samples corresponding to the concatenation of lines spoken by fictional characters.
0 papers · 0 benchmarks
A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects.
0 papers · 0 benchmarks
DOTA2 Games (Dota2 Games Results)
Dota 2 is a popular computer game with two teams of 5 players.
0 papers · 0 benchmarks
We introduce an annotated dataset of five thousand human labeled pareidolic face images, called Faces in Things''.
0 papers · 0 benchmarks
Heel Dataset (Heel Bone X-Ray Dataset)
Heel Bone X-Ray Dataset consists of 3,956 X-ray images of the foot, primarily focused on detecting and classifying heel bone diseases.
0 papers · 0 benchmarks
The International Cardiac Arrest REsearch consortium (I-CARE) Database includes baseline clinical information and continuous electroencephalogram (EEG) and electrocardiogram (ECG) recordings from comatose patients following cardiac arrest.
0 papers · 0 benchmarks
Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift…
0 papers · 0 benchmarks
The thickness and appearance of retinal layers are essential markers for diagnosing and studying eye diseases.
0 papers · 0 benchmarks
Online Shoppers (Online Shoppers Purchasing Intention Dataset)
Of the 12,330 sessions in the dataset, 84.5% (10,422) were negative class samples that did not end with shopping, and the rest (1908) were positive class samples ending with shopping.
0 papers · 0 benchmarks
We propose a new light field image database called “PINet” inheriting the hierarchical structure from WordNet.
0 papers · 0 benchmarks
Sakha-TB (400+400 CXR images for TB diagnosis)
Sakha-TB is a de-identified image dataset of frontal chest X-rays (CXR), collected through collaboration with several medical institutions in the Republic of Sakha (Yakutia, Russia).
0 papers · 0 benchmarks
Semeion (Semeion Handwritten Digit Data Set)
1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.
0 papers · 0 benchmarks
Tornet (Tornado Network)
The Tornado Network (TorNet) dataset is a large, high-resolution benchmark dataset developed to support machine learning research in tornado detection and prediction.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.