Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 5 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 193–240 of 3,130
SimpleQuestions is a large-scale factoid question answering dataset.
120 papers · 2 benchmarks
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
WritingPrompts is a large dataset of 300K human-written stories paired with writing prompts from an online forum.
118 papers · 1 benchmark
FLoRes-200 doubles the existing language coverage of FLoRes-101.
117 papers · 1 benchmark
KILT (Knowledge Intensive Language Tasks) is a benchmark consisting of 11 datasets representing 5 types of tasks: Fact-checking (FEVER), Entity linking (AIDA CoNLL-YAGO, WNED-WIKI, WNED-CWEB), Slot filling (T-Rex, Zero Shot RE), Open…
117 papers · 11 benchmarks
Visual Entailment (VE) consists of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks.
117 papers · 2 benchmarks
The MRQA (Machine Reading for Question Answering) dataset is a dataset for evaluating the generalization capabilities of reading comprehension systems.
116 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
LRS2 (Lip Reading Sentences 2)
The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild.
115 papers · 10 benchmarks
The Stack contains over 3TB of permissively-licensed source code files covering 30 programming languages crawled from GitHub.
115 papers · 0 benchmarks
MCTest is a freely available set of stories and associated questions intended for research on the machine comprehension of text.
114 papers · 2 benchmarks
QASC (Question Answering via Sentence Composition)
QASC is a question-answering dataset with a focus on sentence composition.
114 papers · 0 benchmarks
Visual7W is a large-scale visual question answering (QA) dataset, with object-level groundings and multimodal answers.
112 papers · 1 benchmark
Reading Comprehension with Commonsense Reasoning Dataset (ReCoRD) is a large-scale reading comprehension dataset which requires commonsense reasoning.
111 papers · 1 benchmark
FinQA is a new large-scale dataset with Question-Answering pairs over Financial reports, written by financial experts.
110 papers · 1 benchmark
Dataset Summary Mind2Web is a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website.
110 papers · 1 benchmark
The Cross-lingual Choice of Plausible Alternatives (XCOPA) dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across languages.
110 papers · 1 benchmark
The Evaluation framework of Raganato et al.
109 papers · 3 benchmarks
CommonGen is constructed through a combination of crowdsourced and existing caption corpora, consists of 79k commonsense descriptions over 35k unique concept-sets.
108 papers · 1 benchmark
MGSM (Multilingual Grade School Math)
Multilingual Grade School Math Benchmark (MGSM) is a benchmark of grade-school math problems.
107 papers · 1 benchmark
CORNELL NEWSROOM is a large dataset for training and evaluating summarization systems.
107 papers · 0 benchmarks
VIST (Visual Storytelling)
The Visual Storytelling Dataset (VIST) consists of 210,819 unique photos and 50,000 stories.
107 papers · 2 benchmarks
ReDial (Recommendation Dialogues) is an annotated dataset of dialogues, where users recommend movies to each other.
105 papers · 2 benchmarks
ToolBench is an instruction-tuning dataset for tool use, which is created automatically using ChatGPT.
105 papers · 1 benchmark
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical…
104 papers · 0 benchmarks
ReferIt3D provides two large-scale and complementary visio-linguistic datasets: i) Sr3D, which contains 83.5K template-based utterances leveraging spatial relations among fine-grained object classes to localize a referred object in a…
104 papers · 1 benchmark
GYAFC (Grammarly’s Yahoo Answers Formality Corpus)
Grammarly’s Yahoo Answers Formality Corpus (GYAFC) is the largest dataset for any style containing a total of 110K informal / formal sentence pairs.
103 papers · 2 benchmarks
CosmosQA is a large-scale dataset of 35.6K problems that require commonsense-based reading comprehension, formulated as multiple-choice questions.
102 papers · 0 benchmarks
QASPER is a dataset for question answering on scientific research papers.
102 papers · 1 benchmark
ConvAI2 (Conversational Intelligence Challenge 2)
The ConvAI2 NeurIPS competition aimed at finding approaches to creating high-quality dialogue agents capable of meaningful open domain conversation.
100 papers · 1 benchmark
Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models.
100 papers · 1 benchmark
CLUE (Chinese Language Understanding Evaluation Benchmark)
CLUE is a Chinese Language Understanding Evaluation benchmark.
99 papers · 8 benchmarks
QuALITY (Question Answering with Long Input Texts, Yes!)
QuALITY (Question Answering with Long Input Texts, Yes!) is a multiple-choice question answering dataset for long document comprehension.
98 papers · 1 benchmark
Contains 145k captions for 28k images.
98 papers · 1 benchmark
The ListOps examples are comprised of summary operations on lists of single digit integers, written in prefix notation.
97 papers · 0 benchmarks
FIGER (Fine-Grained Entity Recognition)
The FIGER dataset is an entity recognition dataset where entities are labelled using fine-grained system 112 tags, such as person/doctor, art/writtenwork and building/hotel.
96 papers · 2 benchmarks
RAVEN consists of 1,120,000 images and 70,000 RPM (Raven's Progressive Matrices) problems, equally distributed in 7 distinct figure configurations.
96 papers · 0 benchmarks
Math23K (Math23K for Math Word Problem Solving)
Math23K is a dataset created for math word problem solving, contains 23, 162 Chinese problems crawled from the Internet.
95 papers · 1 benchmark
The CUHK-PEDES dataset is a caption-annotated pedestrian dataset.
93 papers · 3 benchmarks
CBT (Children’s Book Test)
Children’s Book Test (CBT) is designed to measure directly how well language models can exploit wider linguistic context.
92 papers · 1 benchmark
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
JFLEG (JHU FLuency-Extended GUG corpus)
JFLEG is for developing and evaluating grammatical error correction (GEC).
91 papers · 5 benchmarks
PopQA is an open-domain QA dataset with 14k QA pairs with fine-grained Wikidata entity ID, Wikipedia page views, and relationship type information.
91 papers · 1 benchmark
WikiMatrix is a dataset of parallel sentences in the textual content of Wikipedia for all possible language pairs.
91 papers · 0 benchmarks
E2E (End-to-End NLG Challenge)
End-to-End NLG Challenge (E2E) aims to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena.
90 papers · 4 benchmarks
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning.
90 papers · 1 benchmark
The COCO-Text dataset is a dataset for text detection and recognition.
89 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.