Home › Datasets › task › Text Classification
Text Classification datasets
archive 2025-07-28
167 datasets carry the task tag "Text Classification" (the task itself: Text Classification), ordered by the archive's paper count. Page 2 of 4: 48 shown of 167. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Text Classification datasets 49–96 of 167
The Ghostbusters dataset leverages the GPT-3.5-turbo model for generating texts in the domains of creative writing, news, and student essays, providing 2,000 texts in the first two domains and 1,994 in the latter.
13 papers · 1 benchmark
HopeEDI (HopeEDI: A Multilingual Hope Speech Detection Dataset for Equality, Diversity, and Inclusion)
Over the past few years, systems have been developed to control online content and eliminate abusive, offensive or hate speech content.
13 papers · 4 benchmarks
It contains 15K triplets of essay problem statements, student-written, and LLM-generated essays.
13 papers · 0 benchmarks
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
The Overruling dataset is a law dataset corresponding to the task of determining when a sentence is overruling a prior decision.
12 papers · 1 benchmark
The MAGE dataset provides a large set of generated texts using 27 LLMs from seven different groups: OpenAI GPT, LLaMA, GLM130B, FLAN-T5, OPT, BigScience, and EleutherAI.
11 papers · 1 benchmark
Ohsumed includes medical abstracts from the MeSH categories of the year 1991.
11 papers · 2 benchmarks
An expert-annotated word similarity dataset which provides a highly reliable, yet challenging, benchmark for rare word representation techniques.
10 papers · 0 benchmarks
A large-scale curated dataset of over 152 million tweets, growing daily, related to COVID-19 chatter generated from January 1st to April 4th at the time of writing.
10 papers · 0 benchmarks
JGLUE, Japanese General Language Understanding Evaluation, is built to measure the general NLU ability in Japanese.
7 papers · 0 benchmarks
A sentiment analysis Tunisian Arabizi Dataset, collected from social networks, preprocessed for analytical studies and annotated manually by Tunisian native speakers.
7 papers · 0 benchmarks
OMICS (Open Mind Indoor Common Sense)
OMICS is an extensive collection of knowledge for indoor service robots gathered from internet users.
6 papers · 0 benchmarks
TREC-10 (TREC-10 Question Classification)
A question type classification dataset with 6 classes for questions about a person, location, numeric information, etc.
6 papers · 1 benchmark
Benchmark dataset for abstracts and titles of 100,000 ArXiv scientific papers.
6 papers · 1 benchmark
A dataset for evaluating text classification, domain adaptation, and active learning models.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
Taiga is a corpus, where text sources and their meta-information are collected according to popular ML tasks.
5 papers · 0 benchmarks
The topic of Climate Change (CC) has received limited attention in NLP despite its real world urgency.
4 papers · 1 benchmark
The LIAR dataset has been widely followed by fake news detection researchers since its release, and along with a great deal of research, the community has provided a variety of feedback on the dataset to improve it.
4 papers · 1 benchmark
MN-DS (Multilabeled News Dataset)
Multilabeled News Dataset (MN-DS) is a dataset for news classification.
4 papers · 0 benchmarks
The Medical Abstracts dataset contains 14,438 medical abstracts describing 5 different classes of patient conditions, with all of the dataset being annotated.
4 papers · 1 benchmark
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection This is MetaHate: a meta-collection of 36 hate speech datasets from social media comments.
4 papers · 0 benchmarks
MIXSET comprises a total of 3.6k mixtext instances.
4 papers · 1 benchmark
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
A morpho-syntactically annotated Tunisian Arabish Corpus (TArC).
4 papers · 0 benchmarks
This is an entity-level Twitter Sentiment Analysis dataset.
4 papers · 1 benchmark
UIT-VSMEC (Vietnamese Social Media Emotion Corpus)
Emotion recognition is a higher approach or special case of sentiment analysis.
4 papers · 0 benchmarks
FreSaDa is a French satire dataset for cross-domain satire detection, which is composed of 11,570 articles from the news domain.
3 papers · 0 benchmarks
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
We develop a primary dataset based on our task of suicide or depression classification.
3 papers · 0 benchmarks
TCAB (Text Classification Attack Benchmark)
Text Classification Attack Benchmark (TCAB) is a dataset for analyzing, understanding, detecting, and labeling adversarial attacks against text classifiers.
3 papers · 0 benchmarks
Wikipedia Title is a dataset for learning character-level compositionality from the character visual characteristics.
3 papers · 0 benchmarks
Antonio Gulli’s corpus of news articles is a collection of more than 1 million news articles.
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
CIC (Catalonia Independence Corpus)
The dataset is annotated with stance towards one topic, namely, the independence of Catalonia.
2 papers · 3 benchmarks
HLGD (Headline Grouping Dataset)
The Headline Grouping dataset is a binary classification dataset on pairs of news headline.
2 papers · 0 benchmarks
KanHope (Kannada Hope speech dataset)
KanHope is a code mixed hope speech dataset for equality, diversity, and inclusion in Kannada, an under-resourced Dravidian language.
2 papers · 1 benchmark
LatamXIX (19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction)
A novel dataset of 19th-century Latin American press texts, which addresses the lack of specialized corpora for historical and linguistic analysis in this region.
2 papers · 0 benchmarks
Description Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation" (Li et al., 2022).
2 papers · 0 benchmarks
The dataset consists of titles and abstracts from NLP-related papers.
2 papers · 0 benchmarks
The dataset used in the experiments on the paper "Modeling citation worthiness by using attention‑based bidirectional long short‑term memory networks and interpretable models" There are one million sentences in total, and further splitted…
2 papers · 0 benchmarks
RSDD-Time is a dataset of 598 manually annotated self-reported depression diagnosis posts from Reddit that include temporal information about the diagnosis.
2 papers · 0 benchmarks
Articles originating from subreddits with explicitly stated ideologies are categorized into three groups: 72,488 articles in the Liberal class, 79,573 articles in the Conservative class, and 225,083 articles in the Restricted class.
2 papers · 1 benchmark
https://github.com/dialogue-evaluation/RuSentNE-evaluation
2 papers · 0 benchmarks
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
Data set constructed from YouTube comments (72,098 comments posted by 43,859 users on 623 relevant videos to the crisis)
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.