Home › Datasets › task › Text Classification
Text Classification datasets
archive 2025-07-28
167 datasets carry the task tag "Text Classification" (the task itself: Text Classification), ordered by the archive's paper count. Page 3 of 4: 48 shown of 167. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Text Classification datasets 97–144 of 167
Tweets and items from psychological scales for sexism detection with counterfactual examples.
2 papers · 0 benchmarks
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
The data set contains 2500 manually-stance-labeled tweets, 1250 for each candidate (Joe Biden and Donald Trump).
2 papers · 2 benchmarks
AGB-DE is a legal NLP corpus for the automated detection of potentially void clauses in German standard form consumer contracts.
1 paper · 1 benchmark
The AI-GA (Artificial Intelligence Generated Abstracts) dataset is a collection of abstracts and titles, with half of the abstracts being AI-generated and the other half being original.
1 paper · 0 benchmarks
This project contains instructions and codes to reconstruct a dataset for the development and evaluation of forensic tools for detecting machine-generated text in social media.
1 paper · 0 benchmarks
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable.
1 paper · 1 benchmark
The AppealCase dataset is the first large-scale resource specifically designed to support LegalAI research in appellate judgment scenarios.
1 paper · 0 benchmarks
A syllogism is a common form of deductive reasoning that requires precisely two premises and one conclusion.
1 paper · 0 benchmarks
BASIR (BASIR_Budget_Assisted_Sectoral_Impact_Ranking)
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets.
1 paper · 0 benchmarks
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations.
1 paper · 1 benchmark
BanglaBook (Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews)
This repository contains the code, data, and models of the paper titled "BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews" published in the Findings of the Association for Computational Linguistics: ACL…
1 paper · 1 benchmark
BanglaEmotion (BanglaEmotion: A Benchmark Dataset for Bangla Textual Emotion Analysis)
BanglaEmotion is a manually annotated Bangla Emotion corpus, which incorporates the diversity of fine-grained emotion expressions in social-media text.
1 paper · 0 benchmarks
Dataset of 5,591 labeled issue tickets.
1 paper · 0 benchmarks
The dataset contains Tweet IDs along with the location and tweet timestamp.
1 paper · 0 benchmarks
We present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396,209 papers.
1 paper · 0 benchmarks
A dataset of games played in the card game "Cards Against Humanity" (CAH), by human players, derived from the online CAH labs.
1 paper · 0 benchmarks
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
1 paper · 0 benchmarks
Caselaw4 is a dataset of 350k common law judicial decisions from the U.S.
1 paper · 0 benchmarks
Large-scale Chinese legal dataset for judgment prediction.
1 paper · 0 benchmarks
This dataset contains synthetic text data generated to train models for text generation.
1 paper · 0 benchmarks
Dissonance Twitter Dataset is a dataset collected from annotating tweets for dissonance.
1 paper · 0 benchmarks
FMC-MWO2KG (The MWO2KG Failure Mode Classification Dataset)
The Failure Mode Classification dataset released in the paper "MWO2KG and Echidna: Constructing and exploring knowledge graphs from maintenance data" by Stewart et al.
1 paper · 1 benchmark
FTR-18 is a multilingual rumour dataset on football transfer news.
1 paper · 0 benchmarks
FedNLP (FOMC Docs and Speeches)
We collect the various forms of Federal Reserve communications.
1 paper · 0 benchmarks
Enhancing Financial Market Predictions: Causality-Driven Feature Selection This paper introduces FinSen dataset that revolutionizes financial market analysis by integrating economic and financial news articles from 197 countries with stock…
1 paper · 1 benchmark
This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY.
1 paper · 0 benchmarks
The Topic-Based Paragraph Classification in Genocide-Related Court Transcripts (GTC) dataset is the first reference corpus annotated with samples from genocide tribunals in different international criminal courts.
1 paper · 0 benchmarks
A dataset including texts by humans (labeled 0) and then rephrased by ChatGPT (labeled 1), created to train models for machine-generated text detection.
1 paper · 0 benchmarks
Invisible Mobile Keyboard Dataset contains user initial, age, type of mobile devices, size of the screen, time taken for typing each phrase, and annotation of typed phrases with coordinate values of the typed position (x and y points).
1 paper · 0 benchmarks
Dataset Summary New dataset introduced in Parameter-Efficient Legal Domain Adaptation (Li et al., 2022) from the Legal Advice Reddit community (known as "/r/legaldvice"), sourcing the Reddit posts from the Pushshift Reddit dataset.
1 paper · 0 benchmarks
LoT-insts contains over 25k classes whose frequencies are naturally long-tail distributed.
1 paper · 2 benchmarks
M-Phasis (A Feature-Based Corpus of Hate Online)
A corpus of 9k German and French user comments collected from migration-related news articles.
1 paper · 0 benchmarks
The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word…
1 paper · 0 benchmarks
MiST (Modals In Scientific Text) is a dataset containing 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function.
1 paper · 0 benchmarks
Modern Hebrew Sentiment Dataset is a sentiment analysis benchmark for Hebrew, based on 12K social media comments, and provide two instances of these data: in token-based and morpheme-based settings.
1 paper · 0 benchmarks
This is the large version of the MuMiN dataset.
1 paper · 1 benchmark
This is the medium version of the MuMiN dataset.
1 paper · 1 benchmark
This is the small version of the MuMiN dataset.
1 paper · 1 benchmark
This corpus contains data files that were generated as part of the NOVIC paper (see above).
1 paper · 0 benchmarks
A general purpose text categorization dataset (NatCat) from three online resources: Wikipedia, Reddit, and Stack Exchange.
1 paper · 0 benchmarks
Paper Field is built from the Microsoft Academic Graph and maps paper titles to one of 7 fields of study.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE (SILICONE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems specifically designed for spoken language.
1 paper · 1 benchmark
StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic.
1 paper · 1 benchmark
This resource contains 10.5 million paragraphs with associated statement labels, realized as one paragraph per file, one sentence per line.
1 paper · 0 benchmarks
The ShapeIt dataset introduced by Alper et al.
1 paper · 0 benchmarks
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper: Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.