Home › Datasets › task › Text Classification

Text Classification datasets

archive 2025-07-28

167 datasets carry the task tag "Text Classification" (the task itself: Text Classification), ordered by the archive's paper count. Page 1 of 4: 48 shown of 167. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text Classification datasets 1–48 of 167

The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
GLUE (General Language Understanding Evaluation benchmark)
General Language Understanding Evaluation (GLUE) benchmark is a collection of nine natural language understanding tasks, including single-sentence tasks CoLA and SST-2, similarity and paraphrasing tasks MRPC, STS-B and QQP, and natural…
3,197 papers · 13 benchmarks
SST (Stanford Sentiment Treebank)
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language.
2,354 papers · 6 benchmarks
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language.
1,808 papers · 2 benchmarks
The IMDb Movie Reviews dataset is a binary sentiment analysis dataset consisting of 50,000 reviews from the Internet Movie Database (IMDb) labeled as positive or negative.
1,787 papers · 9 benchmarks
AG News (AG’s News Corpus)
AG News (AG’s News Corpus) is a subdataset of AG's corpus of news articles constructed by assembling titles and description fields of articles from the 4 largest classes (“World”, “Sports”, “Business”, “Sci/Tech”) of AG’s Corpus.
969 papers · 9 benchmarks
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples.
701 papers · 5 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in the Wikipedia project.
597 papers · 4 benchmarks
RCV1 (Reuters Corpus Volume 1)
The RCV1 dataset is a benchmark dataset on text categorization.
336 papers · 6 benchmarks
OpenWebText is an open-source recreation of the WebText corpus.
207 papers · 2 benchmarks
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
MTEB (Massive Text Embedding Benchmark)
MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages.
155 papers · 6 benchmarks
e-SNLI is used for various goals, such as obtaining full sentence justifications of a model's decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets.
139 papers · 1 benchmark
CARER (Contextualized Affect Representations for Emotion Recognition)
CARER is an emotion dataset collected through noisy labels, annotated via distant supervision as in (Go et al., 2009).
123 papers · 2 benchmarks
Covers multiple aspects of the issue.
105 papers · 2 benchmarks
CLUE (Chinese Language Understanding Evaluation Benchmark)
CLUE is a Chinese Language Understanding Evaluation benchmark.
99 papers · 8 benchmarks
This dataset is for evaluating the performance of intent classification systems in the presence of "out-of-scope" queries, i.e., queries that do not fall into any of the system-supported intent classes.
87 papers · 5 benchmarks
The Yelp Dataset is a valuable resource for academic research, teaching, and learning.
86 papers · 15 benchmarks
TweetEval introduces an evaluation framework consisting of seven heterogeneous Twitter-specific classification tasks.
84 papers · 1 benchmark
TrecQA (Text Retrieval Conference Question Answering)
Text Retrieval Conference Question Answering (TrecQA) is a dataset created from the TREC-8 (1999) to TREC-13 (2004) Question Answering tracks.
73 papers · 3 benchmarks
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
PearRead is a dataset of scientific peer reviews.
42 papers · 0 benchmarks
This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014.
41 papers · 5 benchmarks
BLURB (Biomedical Language Understanding and Reasoning Benchmark)
BLURB is a collection of resources for biomedical natural language processing.
40 papers · 2 benchmarks
The Yelp Reviews Polarity dataset is obtained from the Yelp Dataset Challenge in 2015 (1,569,264 samples that have review text).
37 papers · 0 benchmarks
Arxiv HEP-TH (high energy physics theory) citation graph is from the e-print arXiv and covers all the citations within a dataset of 27,770 papers with 352,807 edges.
35 papers · 5 benchmarks
MR (MR Movie Reviews)
MR Movie Reviews is a dataset for use in sentiment-analysis experiments.
28 papers · 3 benchmarks
WNUT-2020 Task 2 (WNUT-2020 Task 2: Identification of Informative COVID-19 English Tweets)
Briefly describe the dataset.
28 papers · 1 benchmark
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups.
27 papers · 5 benchmarks
300 news articles annotated with 1,727 bias spans and find evidence that informational bias appears in news articles more frequently than lexical bias.
27 papers · 0 benchmarks
Evidence Inference is a corpus for this task comprising 10,000+ prompts coupled with full-text articles describing RCTs.
27 papers · 0 benchmarks
BioRED is a first-of-its-kind biomedical relation extraction dataset with multiple entity types (e.g.
25 papers · 3 benchmarks
Moral Stories is a crowd-sourced dataset of structured narratives that describe normative and norm-divergent actions taken by individuals to accomplish certain intentions in concrete situations, and their respective consequences.
24 papers · 0 benchmarks
EURLEX57K is a new publicly available legal LMTC dataset, dubbed EURLEX57K, containing 57k English EU legislative documents from the EUR-LEX portal, tagged with ∼4.3k labels (concepts) from the European Vocabulary (EUROVOC).
23 papers · 1 benchmark
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
The Terms of Service dataset is a law dataset corresponding to the task of identifying whether contractual terms are potentially unfair.
21 papers · 1 benchmark
PubMed RCT (PubMed 200k RCT)
PubMed 200k RCT is new dataset based on PubMed for sequential sentence classification.
20 papers · 0 benchmarks
LSHTC is a dataset for large-scale text classification.
18 papers · 0 benchmarks
TuringBench is a benchmark environment that contains : - Benchmark tasks- Turing Test (i.e., human vs.
18 papers · 2 benchmarks
SMHD (Self-reported Mental Health Diagnoses)
A novel large dataset of social media posts from users with one or multiple mental health conditions along with matched control users.
16 papers · 0 benchmarks
The TweepFake dataset consists of 25,572 social media messages posted either by bots or humans on Twitter.
16 papers · 1 benchmark
BeerAdvocate is a dataset that consists of beer reviews from beeradvocate.
15 papers · 1 benchmark
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports.
14 papers · 3 benchmarks
The IndoNLU benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems for Bahasa Indonesia.
14 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.