Home › Datasets › task › Classification

Classification datasets

archive 2025-07-28

222 datasets carry the task tag "Classification" (the task itself: Classification), ordered by the archive's paper count. Page 3 of 5: 48 shown of 222. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Classification datasets 97–144 of 222

NeuroVoz (NeuroVoz: a Castillian Spanish corpus of parkinsonian speech)
The NeuroVoz dataset emerges as a pioneering resource in the field of computational linguistics and biomedical research, specifically designed to enhance the diagnosis and understanding of Parkinson's Disease (PD) through speech analysis.
2 papers · 0 benchmarks
Open Radar Datasets (Open Radar Datasets: Outdoor Moving Object Dataset)
A classification dataset of radar spectrograms in i "ground surveillance" setting recorded with the Open Radar Initiative.
2 papers · 0 benchmarks
Abstract In the view of national security, radar micro-Doppler (m-D) signatures-based recognition of suspicious human activities becomes significant.
2 papers · 2 benchmarks
SKILL-102 (SKILL 102 Lifelong Learning Dataset)
SKILL-102 consists of 102 image classification datasets.
2 papers · 0 benchmarks
The resources for this dataset can be found at https://www.openml.org/d/182 Author: Ashwin Srinivasan, Department of Statistics and Data Modeling, University of Strathclyde Source: UCI - 1993 Please cite: UCI The database consists of the…
2 papers · 0 benchmarks
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
A dataset of 53 complex-valued signal modulation classes.
2 papers · 0 benchmarks
Tiny ImageNet-R is a subset of the ImageNet-R dataset by Hendrycks et al.
2 papers · 0 benchmarks
UK Biobank Brain MRI (UK Biobank Data - Brain MRI)
UK Biobank participants have generously provided a very wide range of information about their health and well-being since recruitment began in 2006.
2 papers · 1 benchmark
Underwater Trash Detection Dataset Overview The Underwater Trash Detection Dataset is a custom-annotated dataset designed to address the challenges of underwater trash detection caused by varying environmental features.
2 papers · 0 benchmarks
news20 (NewsWeeder: learning to filter netnews)
Two datasets featuring binary and multi-class classification.
2 papers · 0 benchmarks
ABUZZ (Citizen-based mosquito monitoring system)
As part of our policy to openly share all data from this project, we have included a downloadable package comprising all acoustic data collected over the course of this work.
1 paper · 0 benchmarks
AUR & UMB dataset (Anticancer Efficacy of Auraptene & Umbelliprenin: In Vitro Viability Dataset)
This dataset contains quantitative data on the anticancer effects of the natural coumarins Auraptene (AUR) and Umbelliprenin (UMB) across 27 studies.
1 paper · 1 benchmark
AjwaOrMedjool (AjwaOrMedjool: a binary balanced dataset to teach machine learning‏)
The dataset contains three subsets: 1- a dataset containing hand-crafted features to classify two types of organic dates (Ajwa or Medjool); 2- a dataset containing tabular data with features created automatically using deep learning to…
1 paper · 0 benchmarks
B-XAIC consists of 50K small molecules represented as graphs and includes 7 graph classification tasks, each with ground truth labels and corresponding explanations.
1 paper · 0 benchmarks
BFN (Backdoored Face-Networks Dataset)
This database is a database of backdoored neural networks intended for face recognition.
1 paper · 0 benchmarks
BFRD (Bengali Fake Review Dataset)
This is a binary dataset used for Bengali fake review detection in the paper "Bengali Fake Reviews: A Benchmark Dataset and Detection System" accepted in Neurocomputing, a journal published by Elsevier.
1 paper · 0 benchmarks
BRISC (BRISC: Annotated Dataset for Brain Tumor Segmentation and Classification)
BRISC is a high-quality, expert-annotated MRI dataset curated for brain tumor segmentation and classification.
1 paper · 1 benchmark
BTS (Building Timeseries Dataset: Empowering Large-Scale Building Analytics)
The Building TimeSeries (BTS) dataset covers three buildings over a three-year period, comprising more than ten thousand timeseries data points with hundreds of unique ontologies.
1 paper · 0 benchmarks
The dataset contains 36000 Bangla data based on Ekman's six basic emotions.
1 paper · 1 benchmark
Original images and images with RUSTICO filters applied Also a csv with classes is included
1 paper · 1 benchmark
CORBEL (Conveyor belt pressure signal dataset))
Dataset included measuring static tension under 2 kg load in different points of the CB and measurements in dynamic conditions.
1 paper · 1 benchmark
CRCDX (TCGA-CRC-DX)
Histological images of colorectal cancer, derived from the TCGA database
1 paper · 0 benchmarks
CUCO Database (A voice and speech corpus of patients who underwent upper airway surgery in pre-and post-operative states)
Many research articles have explored the impact of surgical interventions on voice and speech evaluations, but advances are limited by the lack of publicly accessible datasets.
1 paper · 0 benchmarks
CVE (Common Vulnerabilities and Exposures)
CVE stands for Common Vulnerabilities and Exposures.
1 paper · 0 benchmarks
A dataset of games played in the card game "Cards Against Humanity" (CAH), by human players, derived from the online CAH labs.
1 paper · 0 benchmarks
This dataset derives from Coil100.
1 paper · 0 benchmarks
A large dataset of color names and their respective RGB values stores in CSV.
1 paper · 1 benchmark
This dataset contains synthetic text data generated to train models for text generation.
1 paper · 0 benchmarks
DIGITal (Digitally Generated Numerals)
Digitally Generated Numerals (DIGITal) Description The Digitally Generated Numerals (DIGITal) dataset consists of 100,000 image pairs representing digits from 0 to 9.
1 paper · 0 benchmarks
The dataset comprises motion sensor data of 19 daily and sports activities each performed by 8 subjects in their own style for 5 minutes.
1 paper · 0 benchmarks
DeepGraviLens is a data set of simulated gravitational lenses consisting of images associated with brightness variation time series.
1 paper · 0 benchmarks
DeepParliament is a legal domain Benchmark Dataset that gathers bill documents and metadata and performs various bill status classification tasks.
1 paper · 0 benchmarks
Dhoroni (Dhoroni: A Multi-Perspective Bengali Climate Change and Environmental News Dataset)
Climate change poses critical challenges globally, disproportionately affecting low-income countries that often lack resources and linguistic representation on the international stage.
1 paper · 1 benchmark
DiscoEval (Discourse Evaluation)
Dataset Summary The DiscoEval is an English-language Benchmark that contains a test suite of 7 tasks to evaluate whether sentence representations include semantic information relevant to discourse processing.
1 paper · 0 benchmarks
Dissonance Twitter Dataset is a dataset collected from annotating tweets for dissonance.
1 paper · 0 benchmarks
EuroSAT-C is an open-source data set comprising algorithmically generated corruptions applied to the EuroSAT test set following the concept of ImageNet-C.
1 paper · 0 benchmarks
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by…
1 paper · 0 benchmarks
This dataset was created to test whether it's possible to build a general-purpose detector that can tell real images apart from fake ones generated by convolutional neural networks (CNNs), no matter which model or dataset was used to…
1 paper · 0 benchmarks
GLAMI-1M (A Multilingual Image-Text Fashion Dataset)
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark.
1 paper · 1 benchmark
Gambling Address Dataset is a collection of 10,423 gambling addresses that have transactions with gambling contracts.
1 paper · 0 benchmarks
Gambling Contract Dataset is a collection of 260 gambling smart contracts from decentralized gambling websites, such as Dicether, Degens.
1 paper · 0 benchmarks
- Images for Classification, Segmentation, Object Detection, Upsampling, and Edge LLM - Feature Noise-Augmented Dataset for Semantic Communication Kindly Check: https://huggingface.co/datasets/CQILAB/GenSC-6G
1 paper · 0 benchmarks
This dataset was curated for Search Engine Optimization (SEO) analysis tasks, including categorization and spam detection.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
Dataset introduced by Xifeng Yan et al.
1 paper · 0 benchmarks
HOWS (HOWS-CL-25)
HOWS-CL-25 (Household Objects Within Simulation dataset for Continual Learning) is a synthetic dataset especially designed for object classification on mobile robots operating in a changing environment (like a household), where it is…
1 paper · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.