Home › Datasets › task › Classification

Classification datasets

archive 2025-07-28

222 datasets carry the task tag "Classification" (the task itself: Classification), ordered by the archive's paper count. Page 4 of 5: 48 shown of 222. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Classification datasets 145–192 of 222

HRPlanesV2 (HRPlanesv2 - High Resolution Satellite Imagery for Aircraft Detection)
The HRPlanesv2 dataset contains 2120 VHR Google Earth images.
1 paper · 0 benchmarks
This dataset was presented as part of the ICLR 2023 paper 𝘈 𝘧𝘳𝘢𝘮𝘦𝘸𝘰𝘳𝘬 𝘧𝘰𝘳 𝘣𝘦𝘯𝘤𝘩𝘮𝘢𝘳𝘬𝘪𝘯𝘨 𝘊𝘭𝘢𝘴𝘴-𝘰𝘶𝘵-𝘰𝘧-𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 𝘥𝘦𝘵𝘦𝘤𝘵𝘪𝘰𝘯 𝘢𝘯𝘥 𝘪𝘵𝘴 𝘢𝘱𝘱𝘭𝘪𝘤𝘢𝘵𝘪𝘰𝘯 𝘵𝘰 𝘐𝘮𝘢𝘨𝘦𝘕𝘦𝘵.
1 paper · 1 benchmark
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs.
1 paper · 0 benchmarks
This is a real-world industrial benchmark dataset from a major medical device manufacturer for the prediction of customer escalations.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
JAMBO (A Multi-Annotator Image Dataset for Benthic Habitat Classification)
The JAMBO dataset contains 3290 underwater images of the seabed captured by an ROV in temperate waters in the Jammer Bay area off the North West coast of Jutland, Denmark.
1 paper · 0 benchmarks
Context The Kepler Space Observatory is a NASA-build satellite that was launched in 2009.
1 paper · 1 benchmark
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
This data set comprises 22 fundus images with their corresponding manual annotations for the blood vessels, separated as arteries and veins.
1 paper · 2 benchmarks
Liver-US (Liver Ultrasound Dataset for Medical Image Classification)
The Liver-US dataset is a comprehensive collection of high-quality ultrasound images of the liver, including both normal and abnormal cases.
1 paper · 1 benchmark
LoRA-WiSE (LoRA Weight Size Evaluation)
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models LoRA-WiSE spans various dataset sizes, backbones, ranks, and…
1 paper · 0 benchmarks
MARIO (Monitoring Age-related Macular Degeneration Progression In Optical Coherence Tomography)
MICCAI Challenge 2024
1 paper · 0 benchmarks
MVTec-FS (MVTec few-shot detection and classfication dataset)
The MVTec-FS dataset is a refined version of the MVTec AD dataset, designed for few-shot learning research.
1 paper · 0 benchmarks
MalVis (MalVis: A Large-Scale Android Malware Visualization Dataset and Framework for Improved Classification)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MapReader Data (in GeoHumanities workshop, SIGSPATIAL 2022)
MapReader in GeoHumanities workshop (SIGSPATIAL 2022): Gold standards and outputs Refer to: https://github.com/Living-with-machines/MapReader/wiki/GeoHumanities-workshop-in-SIGSPATIAL-2022
1 paper · 0 benchmarks
A large-scale reference dataset for bioacoustics.
1 paper · 1 benchmark
MiST (Modals In Scientific Text) is a dataset containing 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function.
1 paper · 0 benchmarks
MixedWM38 Dataset(WaferMap) has more than 38000 wafer maps, including 1 normal pattern, 8 single defect patterns, and 29 mixed defect patterns, a total of 38 defect patterns.
1 paper · 2 benchmarks
Dataset can be used by anyone who is interested to perform morphological classification of galaxies.
1 paper · 0 benchmarks
MuReD Dataset (Multi-Label Retinal Diseases Dataset)
Early detection of retinal diseases is one of the most important means of preventing partial or permanent blindness in patients.
1 paper · 1 benchmark
NBA_Box_Scores_Odds (NBA Team-Level Box Score Statistics (2015-2019), Historical Win Percentages (2014-2018) and Betting Odds (2018/2019))
Dataset Description: NBA Team Statistics, Historical Performance & Betting Odds (2015-2019) Overview This dataset contains team-level box score statistics, historical win percentages, and closing betting odds for NBA games from 2015 to…
1 paper · 0 benchmarks
Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities, including images, shapes, and scenes.
1 paper · 0 benchmarks
Hand-labelled dataset of crop and non-crop labels distributed throughout Nigeria with respective hd5f data arrays.
1 paper · 0 benchmarks
Niramai Oncho Dataset (Niramai Onchocerciasis/RiverBlindness Dataset)
Onchocerciasis is causing blindness in over half a million people in the world today.
1 paper · 0 benchmarks
Dataset composed of two main parts 1.
1 paper · 0 benchmarks
Monitoring and evaluating of driving behavior is the main goal of this paper that encourage us to develop a new system based on Inertial Measurement Unit (IMU) sensors of smartphones.
1 paper · 0 benchmarks
PASSION dataset (PASSION derm 2024 dataset)
Overview PASSION derm is a pioneering initiative dedicated to closing the diversity gap in dermatology datasets.
1 paper · 0 benchmarks
Physical concept understanding benchmark.
1 paper · 0 benchmarks
In this paper, we propose RFUAV as a new benchmark dataset for radio-frequency based (RF-based) unmanned aerial vehicle (UAV) identification and address the following challenges: Firstly, many existing datasets feature a restricted variety…
1 paper · 0 benchmarks
RGZ EMU: Semantic Taxonomy (Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy)
The data used in - "Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy" (Bowles et al.
1 paper · 0 benchmarks
Raw-Microscopy: 940 raw bright-field microscopy images of human blood smear slides for leukocyte classification (microscopy/images/rawscale100) with corresponding labels (microscopy/labels).
1 paper · 0 benchmarks
Dataset with articles posted in the r/Liberal and r/Conservative subreddits.
1 paper · 1 benchmark
This dataset was acquired in a retrospective study from a cohort of pediatric patients admitted with abdominal pain to Children’s Hospital St.
1 paper · 0 benchmarks
Dataset contains light curves of 6 rocket body types from Mini Mega Tortora database (MMT)[^1].
1 paper · 0 benchmarks
SF-MASK (Small Face MASK)
SF-MASK is a collection made from 20k low-resolution images exported from diverse and heterogeneous datasets, ranging from 7 x 7 to 64 x 64 pixel resolution.
1 paper · 0 benchmarks
SHADR (sythetic SDoH Human Annotated Demographic Robustness dataset (SHADR))
SDoH Human Annotated Demoographic Robustness (SHADR) Dataset Overview The Social determinants of health (SDoH) play a pivotal role in determining patient outcomes.
1 paper · 0 benchmarks
SHD - Adding (Spiking Heidelberg Digits - Adding)
This dataset is based on the Spiking Heidelberg Digits (SHD) dataset.
1 paper · 1 benchmark
SPOT-10 (Animal Pattern Benchmark Dataset for Machine Learning Algorithms)
The SPOTS-10 dataset is an extensive collection of grayscale images showcasing diverse patterns found in ten animal species.
1 paper · 1 benchmark
StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic.
1 paper · 1 benchmark
The Satellite dataset forms a practical VFL scenario for location identification based on satellite imagery.
1 paper · 0 benchmarks
SimGas (Computer Simulated Gas Leakage Segmentation)
This dataset consists of computer-generated images for gas leakage segmentation.
1 paper · 2 benchmarks
Simulated pulse Doppler radar signatures for four classes of helicopter-like targets.
1 paper · 0 benchmarks
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics | Category | Samples | Description |…
1 paper · 0 benchmarks
arxiv : https://arxiv.org/abs/2304.11708 Accepted at 29th International Congress on Sound and Vibration (ICSV29).
1 paper · 1 benchmark
Classifying Email as Spam or Non-Spam.
1 paper · 0 benchmarks
Spanish Corpus XIX (19th Century Spanish Corpus)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
TACM12K (Table-ACM12K)
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset.
1 paper · 1 benchmark
TCB-DS (Toxigenic Cyanobacteria Dataset)
The TCB-DS dataset is a specialized collection of microscopic images focusing on the automatic recognition of cyanobacteria genera.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.