Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 9 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 385–432 of 3,130
GeoQA (Geometric Question Answering)
GeoQA is a dataset for automatic geometric problem solving containing 5,010 geometric problems with corresponding annotated programs, which illustrate the solving process of the given problems Compared with another publicly available…
46 papers · 1 benchmark
Legal General Language Understanding Evaluation (LexGLUE) benchmark is a collection of datasets for evaluating model performance across a diverse set of legal NLU tasks in a standardized way.
46 papers · 1 benchmark
MAMS (Multi Aspect Multi-Sentiment)
MAMS is a challenge dataset for aspect-based sentiment analysis (ABSA), in which each sentences contain at least two aspects with different sentiment polarities.
46 papers · 1 benchmark
UDC (Ubuntu Dialogue Corpus)
Ubuntu Dialogue Corpus (UDC) is a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words.
46 papers · 8 benchmarks
CoSQL (Conversational Text-to-SQL Challenge)
CoSQL is a corpus for building cross-domain, general-purpose database (DB) querying dialogue systems.
45 papers · 1 benchmark
DART is a large dataset for open-domain structured data record to text generation.
45 papers · 3 benchmarks
GEM (Generation, Evaluation, and Metrics)
Generation, Evaluation, and Metrics (GEM) is a benchmark environment for Natural Language Generation with a focus on its Evaluation, both through human annotations and automated Metrics.
45 papers · 1 benchmark
A new large-scale geometry problem-solving dataset - 3,002 multi-choice geometry problems - dense annotations in formal language for the diagrams and text - 27,213 annotated diagram logic forms (literals) - 6,293 annotated text logic forms…
45 papers · 1 benchmark
A new dataset of 1,001 human-human dialogs for movie recommendation with measures for successful recommendations.
45 papers · 0 benchmarks
MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
RSTPReid (Real Scenario Text-based Person Re-identification)
RSTPReid contains 20505 images of 4,101 persons from 15 cameras.
45 papers · 2 benchmarks
EmotionLines contains a total of 29245 labeled utterances from 2000 dialogues.
44 papers · 1 benchmark
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
OCNLI (Original Chinese Natural Language Inference)
OCNLI stands for Original Chinese Natural Language Inference.
44 papers · 0 benchmarks
The SCUT-CTW1500 dataset contains 1,500 images: 1,000 for training and 500 for testing.
44 papers · 3 benchmarks
TurkCorpus, a dataset with 2,359 original sentences from English Wikipedia, each with 8 manual reference simplifications.
44 papers · 1 benchmark
For understanding multimodal language used in expressing humor.
44 papers · 0 benchmarks
Dataset contains 33,010 molecule-description pairs split into 80\%/10\%/10\% train/val/test splits.
43 papers · 4 benchmarks
QA-SRL was proposed as an open schema for semantic roles, in which the relation between an argument and a predicate is expressed as a natural-language question containing the predicate (“Where was someone educated?”) whose answer is the…
43 papers · 0 benchmarks
ShARC (Shaping Answers with Rules through Conversation)
ShARC is a Conversational Question Answering dataset focussing on question answering from texts containing rules.
43 papers · 0 benchmarks
WebQA, is a new benchmark for multimodal multihop reasoning in which systems are presented with the same style of data as humans when searching the web: Snippets and Images.
43 papers · 0 benchmarks
gRefCOCO is the first large-scale Generalized Referring Expression Segmentation dataset that contains multi-target, no-target, and single-target expressions.
43 papers · 2 benchmarks
BUCC (Building and Using Comparable Corpora)
The BUCC mining task is a shared task on parallel sentence extraction from two monolingual corpora with a subset of them assumed to be parallel, and that has been available since 2016.
42 papers · 4 benchmarks
DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie.
42 papers · 1 benchmark
We release E-commerce Dialogue Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
42 papers · 1 benchmark
EmoContext consists of three-turn English Tweets.
42 papers · 1 benchmark
English Web Treebank is a dataset containing 254,830 word-level tokens and 16,624 sentence-level tokens of webtext in 1174 files annotated for sentence- and word-level tokenization, part-of-speech, and syntactic structure.
42 papers · 0 benchmarks
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
PearRead is a dataset of scientific peer reviews.
42 papers · 0 benchmarks
SCROLLS (Standardized CompaRison Over Long Language Sequences)
SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts.
42 papers · 1 benchmark
This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014.
41 papers · 5 benchmarks
A creative writing task where the input is 4 random sentences and the output should be a coherent passage with 4 paragraphs that end in the 4 input sentences respectively.
41 papers · 0 benchmarks
The Query-based Video Highlights (QVHighlights) dataset is a dataset for detecting customized moments and highlights from videos given natural language (NL).
41 papers · 4 benchmarks
A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research.
41 papers · 0 benchmarks
BLURB (Biomedical Language Understanding and Reasoning Benchmark)
BLURB is a collection of resources for biomedical natural language processing.
40 papers · 2 benchmarks
Contains around 200K dialogs with a total of 1.6M turns.
40 papers · 0 benchmarks
RedCaps is a large-scale dataset of 12M image-text pairs collected from Reddit.
40 papers · 0 benchmarks
Samanantar is the largest publicly available parallel corpora collection for Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu.
40 papers · 0 benchmarks
SciCite is a dataset of citation intents that addresses multiple scientific domains and is more than five times larger than ACL-ARC.
40 papers · 3 benchmarks
We propose the first question-answering dataset driven by STEM theorems.
40 papers · 1 benchmark
ALCE (Automatic LLMs' Citation Evaluation)
ALCE is a benchmark for Automatic LLMs' Citation Evaluation.
39 papers · 0 benchmarks
Break is a question understanding dataset, aimed at training models to reason over complex questions.
39 papers · 0 benchmarks
BookSum is a collection of datasets for long-form narrative summarization.
39 papers · 2 benchmarks
The Memetracker corpus contains articles from mainstream media and blogs from August 1 to October 31, 2008 with about 1 million documents per day.
39 papers · 1 benchmark
WikiMovies is a dataset for question answering for movies content.
39 papers · 0 benchmarks
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
38 papers · 0 benchmarks
BIOSSES (Biomedical Semantic Similarity Estimation System)
The BIOSSES data set comprises total 100 sentence pairs all of which were selected from the "TAC2 Biomedical Summarization Track Training Data Set" .
38 papers · 2 benchmarks
CUAD (Contract Understanding Atticus Dataset)
Contract Understanding Atticus Dataset (CUAD) is a dataset for legal contract review.
38 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.