Home › Datasets › task › Text Classification
Text Classification datasets
archive 2025-07-28
167 datasets carry the task tag "Text Classification" (the task itself: Text Classification), ordered by the archive's paper count. Page 4 of 4: 23 shown of 167. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Text Classification datasets 145–167 of 167
ShortPersianEmo is a new data set for emotion recognition in Persian short texts.
1 paper · 1 benchmark
SmokEng is a dataset of 3144 tweets, which are selected based on the presence of colloquial slang related to smoking and analyze it based on the semantics of the tweet.
1 paper · 0 benchmarks
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”.
1 paper · 0 benchmarks
This is not a Dataset (This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models)
We introduce a large semi-automatically generated dataset of ~400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms that we use to…
1 paper · 1 benchmark
6000 French user reviews from three applications on Google Play (Garmin Connect, Huawei Health, Samsung Health) are labelled manually.
1 paper · 0 benchmarks
The Trilemma of Truth is a multiclass probing dataset for evaluating the veracity-tracking mechanism of large language models.
1 paper · 0 benchmarks
TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable.
1 paper · 0 benchmarks
This dataset for abusive content detection in Twitter consists of two sets of annotations for the same set of tweets, one where the human annotators had access to the tweet's content and one where they didn't know the context.
1 paper · 0 benchmarks
Education is increasingly data-driven, and the ability to analyse and adapt educational materials quickly and effectively is important for keeping materials contemporary and interesting.
1 paper · 1 benchmark
This dataset includes User Story (or Issue) text descriptions, User Story titles, and Story Points from 33 software development projects, comprising a total of 20,479 User Stories (or issues) extracted from GitLab repositories, amounting…
1 paper · 0 benchmarks
Wiki-Reliability is the first dataset of English Wikipedia articles annotated with a wide set of content reliability issues.
1 paper · 0 benchmarks
Wiki-en is an annotated English dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
Wiki-zh is an annotated Chinese dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
This is a dataset of scientific documents derived from arXiv.
1 paper · 0 benchmarks
iLur News Texts is a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic.
0 papers · 0 benchmarks
Eduge (Eduge news classification dataset)
Eduge news classification dataset provided by Bolorsoft LLC.
0 papers · 0 benchmarks
A dataset specifically tailored to the biotech news sector, aiming to transcend the limitations of existing benchmarks.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 20,000+ original Number plate images captured and crowdsourced from over 700+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals…
0 papers · 0 benchmarks
MNAD (Moroccan News Articles Dataset)
About the MNAD Dataset The MNAD corpus is a collection of over 1 million Moroccan news articles written in modern Arabic language.
0 papers · 0 benchmarks
RuADReCT (The Russian Adverse Drug Reaction Corpus of Tweets)
Created as part of the Social Media Mining for Health Applications (#SMM4H '20) shared tasks, this dataset consists of 9515 tweets describing health issues.
0 papers · 0 benchmarks
Este conjunto de datos consiste en comentarios de publicaciones del MINSA (Perú) en Facebook sobre la vacuna contra el VPH entre los años 2019 y 2020.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.