Home › Datasets › task › Hierarchical Multi-label Classification
Hierarchical Multi-label Classification datasets
archive 2025-07-28
15 datasets carry the task tag "Hierarchical Multi-label Classification" (the task itself: Hierarchical Multi-label Classification), ordered by the archive's paper count. Page 1 of 1: 15 shown of 15. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Hierarchical Multi-label Classification datasets 1–15 of 15
RCV1 (Reuters Corpus Volume 1)
The RCV1 dataset is a benchmark dataset on text categorization.
336 papers · 6 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
EURLEX57K is a new publicly available legal LMTC dataset, dubbed EURLEX57K, containing 57k English EU legislative documents from the EUR-LEX portal, tagged with ∼4.3k labels (concepts) from the European Vocabulary (EUROVOC).
23 papers · 1 benchmark
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
Hierarchical multi-label classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
Hierarchical-multilabel classification dataset for functional genomics
2 papers · 1 benchmark
The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.