Home › Datasets › task › Document Classification
Document Classification datasets
archive 2025-07-28
18 datasets carry the task tag "Document Classification" (the task itself: Document Classification), ordered by the archive's paper count. Page 1 of 1: 18 shown of 18. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Document Classification datasets 1–18 of 18
The Cora dataset consists of 2708 scientific publications classified into one of seven classes.
602 papers · 18 benchmarks
The MPQA Opinion Corpus contains 535 news articles from a wide variety of news sources manually annotated for opinions and other private states (i.e., beliefs, emotions, sentiments, speculations, etc.).
313 papers · 3 benchmarks
IMDB-MULTI is a relational dataset that consists of a network of 1000 actors or actresses who played roles in movies in IMDB.
243 papers · 3 benchmarks
The Yelp Dataset is a valuable resource for academic research, teaching, and learning.
86 papers · 15 benchmarks
The Reuters-21578 dataset is a collection of documents with news articles.
66 papers · 5 benchmarks
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
SciDocs evaluation framework consists of a suite of evaluation tasks designed for document-level tasks.
57 papers · 2 benchmarks
HOC (Hallmarks of Cancer)
The Hallmarks of Cancer (HOC) corpus consists of 1852 PubMed publication abstracts manually annotated by experts according to the Hallmarks of Cancer taxonomy.
37 papers · 1 benchmark
MultiEURLEX is a multilingual dataset for topic classification of legal documents.
11 papers · 0 benchmarks
LUN is used for unreliable news source classification, this dataset includes 17,250 articles from satire, propaganda, and hoaxe.
8 papers · 1 benchmark
Introduces three datasets of expressing hate, commonly used topics, and opinions for hate speech detection, document classification, and sentiment analysis, respectively.
6 papers · 0 benchmarks
Hyperpartisan News Detection was a dataset created for PAN @ SemEval 2019 Task 4.
3 papers · 1 benchmark
Wikipedia Title is a dataset for learning character-level compositionality from the character visual characteristics.
3 papers · 0 benchmarks
RTC is a benchmark corpus of social media comments sampled over three years.
2 papers · 0 benchmarks
MeSHup (A Corpus for Full Text Biomedical Document Indexing)
Contains 1,342,667 full text articles in English, together with the associated MeSH labels and metadata, authors, and publication venues that are collected from the MEDLINE database.
1 paper · 0 benchmarks
RVL-CDIPMP is our first contribution to retrieve the original documents of the IIT-CDIP test collection which were used to create RVL-CDIP.
1 paper · 0 benchmarks
RVL-CDIPMP-N can serve its original goal as a covariate shift test set, now for multi-page document classification.
1 paper · 0 benchmarks
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.