Home › Datasets › task › Multi-class Classification

Multi-class Classification datasets

archive 2025-07-28

14 datasets carry the task tag "Multi-class Classification" (the task itself: Multi-class Classification), ordered by the archive's paper count. Page 1 of 1: 14 shown of 14. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Multi-class Classification datasets 1–14 of 14

SMM4H (Social Media Mining for Health Shared Task)
Social Media Mining for Health (SMM4H) Shared Task is a massive data source for biomedical and public health applications.
41 papers · 0 benchmarks
TuringBench is a benchmark environment that contains : - Benchmark tasks- Turing Test (i.e., human vs.
18 papers · 2 benchmarks
The twitter emoji dataset obtained from CodaLab comprises of 50 thousand tweets along with the associated emoji label.
4 papers · 0 benchmarks
The TII-SSRC-23 dataset offers a comprehensive collection of network traffic patterns, meticulously compiled to support the development and research of Intrusion Detection Systems (IDS).
2 papers · 3 benchmarks
The AppealCase dataset is the first large-scale resource specifically designed to support LegalAI research in appellate judgment scenarios.
1 paper · 0 benchmarks
CPCXR (COVID-19 Posteroanterior Chest X-Ray fused)
The COVID-19 Posteroanterior Chest X-Ray fused (CPCXR) dataset is generated by the fusion of three publicly available datasets: COVID-19 cxr image, Radiological Society of North America (RSNA), and U.S.
1 paper · 0 benchmarks
CVE (Common Vulnerabilities and Exposures)
CVE stands for Common Vulnerabilities and Exposures.
1 paper · 0 benchmarks
DeepParliament is a legal domain Benchmark Dataset that gathers bill documents and metadata and performs various bill status classification tasks.
1 paper · 0 benchmarks
Two news datasets (KINNEWS and KIRNEWS) for multi-class classification of news articles in Kinyarwanda and Kirundi, two low-resource African languages.
1 paper · 0 benchmarks
The ROAD dataset is made up of observations from the Low Frequency Array (LOFAR) telescope.
1 paper · 0 benchmarks
SF-MASK (Small Face MASK)
SF-MASK is a collection made from 20k low-resolution images exported from diverse and heterogeneous datasets, ranging from 7 x 7 to 64 x 64 pixel resolution.
1 paper · 0 benchmarks
SmokEng is a dataset of 3144 tweets, which are selected based on the presence of colloquial slang related to smoking and analyze it based on the semantics of the tweet.
1 paper · 0 benchmarks
This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware.
0 papers · 0 benchmarks
Eduge (Eduge news classification dataset)
Eduge news classification dataset provided by Bolorsoft LLC.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.