Home › Datasets › task › Multi-Label Text Classification
Multi-Label Text Classification datasets
archive 2025-07-28
13 datasets carry the task tag "Multi-Label Text Classification" (the task itself: Multi-Label Text Classification), ordered by the archive's paper count. Page 1 of 1: 13 shown of 13. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Multi-Label Text Classification datasets 1–13 of 13
MIMIC-III (The Medical Information Mart for Intensive Care III)
The Medical Information Mart for Intensive Care III (MIMIC-III) dataset is a large, de-identified and publicly-available collection of medical records.
1,041 papers · 8 benchmarks
RCV1 (Reuters Corpus Volume 1)
The RCV1 dataset is a benchmark dataset on text categorization.
336 papers · 6 benchmarks
The Reuters-21578 dataset is a collection of documents with news articles.
66 papers · 5 benchmarks
The Slashdot dataset is a relational dataset obtained from Slashdot.
57 papers · 2 benchmarks
ContractNLI is a dataset for document-level natural language inference (NLI) on contracts whose goal is to automate/support a time-consuming procedure of contract review.
31 papers · 0 benchmarks
EURLEX57K is a new publicly available legal LMTC dataset, dubbed EURLEX57K, containing 57k English EU legislative documents from the EUR-LEX portal, tagged with ∼4.3k labels (concepts) from the European Vocabulary (EUROVOC).
23 papers · 1 benchmark
The objective in extreme multi-label classification is to learn feature architectures and classifiers that can automatically tag a data point with the most relevant subset of labels from an extremely large label set.
18 papers · 0 benchmarks
The dataset offers tag and mask annotations for image-text pairs from the CC3M validation set.
5 papers · 2 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
Antibody Watch is a dataset of text snippets extracted from over 2000 PubMed articles with annotations denoting specificity of antibodies.
1 paper · 0 benchmarks
This data is for the Mis2-KDD 2021 under review paper: Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People’s Republic of China We present our dataset that focuses on propaganda techniques in Mandarin…
1 paper · 1 benchmark
1.9K Korean Online Hate Speech Comments for Multilabel Classification (Annotated by Three Independent Labelers per Data)
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.