Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 2 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 49–96 of 3,130

DROP (Discrete Reasoning Over Paragraphs)
Discrete Reasoning Over Paragraphs DROP is a crowdsourced, adversarially-created, 96k-question benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over…
382 papers · 3 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images.
366 papers · 7 benchmarks
SVAMP (Simple Variations on Arithmetic Math word Problems)
A challenge set for elementary-level Math Word Problems (MWP).
362 papers · 2 benchmarks
WSC (Winograd Schema Challenge)
The Winograd Schema Challenge was introduced both as an alternative to the Turing Test and as a test of a system’s ability to do commonsense reasoning.
361 papers · 2 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
XNLI (Cross-lingual Natural Language Inference)
The Cross-lingual Natural Language Inference (XNLI) corpus is the extension of the Multi-Genre NLI (MultiNLI) corpus to 15 languages.
349 papers · 7 benchmarks
SICK (Sentences Involving Compositional Knowledge)
The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics.
348 papers · 5 benchmarks
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
339 papers · 2 benchmarks
ScienceQA (Science Question Answering)
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
RCV1 (Reuters Corpus Volume 1)
The RCV1 dataset is a benchmark dataset on text categorization.
336 papers · 6 benchmarks
COPA (Choice of Plausible Alternatives)
The Choice Of Plausible Alternatives (COPA) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
329 papers · 1 benchmark
MultiWOZ (Multi-domain Wizard-of-Oz)
The Multi-domain Wizard-of-Oz (MultiWOZ) dataset is a large-scale human-human conversational corpus spanning over seven domains, containing 8438 multi-turn dialogues, with each dialogue averaging 14 turns.
328 papers · 8 benchmarks
MSVD (Microsoft Research Video Description Corpus)
The Microsoft Research Video Description Corpus (MSVD) dataset consists of about 120K sentences collected during the summer of 2010.
327 papers · 3 benchmarks
LJSpeech (The LJ Speech Dataset)
This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books.
323 papers · 2 benchmarks
MPQA Opinion Corpus (Multi-Perspective Question Answering)
The MPQA Opinion Corpus contains 535 news articles from a wide variety of news sources manually annotated for opinions and other private states (i.e., beliefs, emotions, sentiments, speculations, etc.).
313 papers · 3 benchmarks
BEIR (Benchmarking IR)
BEIR (Benchmarking IR) is a heterogeneous benchmark containing different information retrieval (IR) tasks.
311 papers · 10 benchmarks
The CodeSearchNet Corpus is a large dataset of functions with associated documentation written in Go, Java, JavaScript, PHP, Python, and Ruby from open source projects on GitHub.
308 papers · 12 benchmarks
The LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects) benchmark is an open-ended cloze task which consists of about 10,000 passages from BooksCorpus where a missing target word is predicted in the last sentence of each…
293 papers · 1 benchmark
StrategyQA is a question answering benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy.
291 papers · 1 benchmark
DocVQA consists of 50,000 questions defined on 12,000+ document images.
290 papers · 3 benchmarks
MELD (Multimodal EmotionLines Dataset)
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
289 papers · 3 benchmarks
WMT 2014 is a collection of datasets used in shared tasks of the Ninth Workshop on Statistical Machine Translation.
288 papers · 9 benchmarks
ANLI (Adversarial NLI)
The Adversarial Natural Language Inference (ANLI, Nie et al.) is a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.
287 papers · 3 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
CoQA (Conversational Question Answering Challenge)
CoQA is a large-scale dataset for building Conversational Question Answering systems.
281 papers · 2 benchmarks
AudioCaps is a dataset of sounds with event descriptions that was introduced for the task of audio captioning, with sounds sourced from the AudioSet dataset.
279 papers · 6 benchmarks
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.
276 papers · 3 benchmarks
The NewsQA dataset is a crowd-sourced machine reading comprehension dataset of 120,000 question-answer pairs.
272 papers · 1 benchmark
WikiSQL consists of a corpus of 87,726 hand-annotated SQL query and natural language question pairs.
267 papers · 4 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
VizWiz (VizWiz-VQA)
The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question.
260 papers · 7 benchmarks
LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, prepared by Heiga Zen with the assistance of Google Speech and Google Brain team members.
257 papers · 1 benchmark
WebVid contains 10 million video clips with captions, sourced from the web.
257 papers · 1 benchmark
SNIPS (SNIPS Natural Language Understanding benchmark)
The SNIPS Natural Language Understanding benchmark is a dataset of over 16,000 crowdsourced queries distributed among 7 user intents of various complexity: SearchCreativeWork (e.g.
256 papers · 6 benchmarks
The ActivityNet Captions dataset is built on ActivityNet v1.3 which includes 20k YouTube untrimmed videos with 100k caption annotations.
255 papers · 6 benchmarks
OntoNotes 5.0 is a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information…
254 papers · 12 benchmarks
The ICDAR 2013 dataset consists of 229 training images and 233 testing images, with word-level annotations provided.
246 papers · 3 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
The WebQuestions dataset is a question answering dataset using Freebase as the knowledge base and contains 6,642 question-answer pairs.
241 papers · 4 benchmarks
MIMIC-CXR from Massachusetts Institute of Technology presents 371,920 chest X-rays associated with 227,943 imaging studies from 65,079 patients.
240 papers · 3 benchmarks
Charades-STA is a new dataset built on top of Charades by adding sentence temporal annotations.
236 papers · 4 benchmarks
DiDeMo (Distinct Describable Moments)
The Distinct Describable Moments (DiDeMo) dataset is one of the largest and most diverse datasets for the temporal localization of events in videos given natural language descriptions.
216 papers · 3 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
OpenWebText is an open-source recreation of the WebText corpus.
207 papers · 2 benchmarks
The NarrativeQA dataset includes a list of documents with Wikipedia summaries, links to full stories, and questions and answers.
206 papers · 1 benchmark
WiC (Words in Context)
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.