Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 7 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 289–336 of 3,130

T2I-CompBench is a comprehensive benchmark for open-world compositional text-to-image generation, consisting of 6,000 compositional textual prompts from 3 categories (attribute binding, object relationships, and complex compositions) and 6…
67 papers · 1 benchmark
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
This project contains natural language data for human-robot interaction in home domain which we collected and annotated for evaluating NLU Services/platforms.
66 papers · 3 benchmarks
Natural-Instructions is a dataset of 61 distinct tasks, their human-authored instructions and 193k task instances.
66 papers · 0 benchmarks
One Billion Word Benchmark (One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling)
Text corpus with almost one billion words of training data for statistical language modeling benchmarking.
66 papers · 0 benchmarks
ParaCrawl v.7.1 is a parallel dataset with 41 language pairs primarily aligned with English (39 out of 41) and mined using the parallel-data-crawling tool Bitextor which includes downloading documents, preprocessing and normalization,…
66 papers · 0 benchmarks
WikiHop is a multi-hop question-answering dataset.
66 papers · 2 benchmarks
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
DuReader is a large-scale open-domain Chinese machine reading comprehension dataset.
65 papers · 1 benchmark
This dataset consists of (human-written) NBA basketball game summaries aligned with their corresponding box- and line-scores.
65 papers · 5 benchmarks
AIDA CoNLL-YAGO contains assignments of entities to the mentions of named entities annotated for the original CoNLL 2003 entity recognition task.
64 papers · 0 benchmarks
AQUA-RAT (Algebra Question Answering with Rationales)
Algebra Question Answering with Rationales (AQUA-RAT) is a dataset that contains algebraic word problems with rationales.
64 papers · 0 benchmarks
The EmpatheticDialogues dataset is a large-scale multi-turn empathetic dialogue dataset collected on the Amazon Mechanical Turk, containing 24,850 one-to-one open-domain conversations.
64 papers · 2 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
SLAKE is an English-Chinese bilingual dataset consisting of 642 images and 14,028 question-answer pairs for training and testing Med-VQA systems.
63 papers · 0 benchmarks
Sentence Compression is a dataset where the syntactic trees of the compressions are subtrees of their uncompressed counterparts, and hence where supervised systems which require a structural alignment between the input and output can be…
63 papers · 0 benchmarks
ComplexWebQuestions is a dataset for answering complex questions that require reasoning over multiple web snippets.
62 papers · 2 benchmarks
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.
62 papers · 2 benchmarks
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
62 papers · 0 benchmarks
CIRR (Compose Image Retrieval on Real-life images)
Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.
61 papers · 3 benchmarks
FigureQA is a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images.
61 papers · 1 benchmark
Game of 24 is a mathematical reasoning challenge, where the goal is to use 4 numbers and basic arithmetic operations (+-/) to obtain 24.
61 papers · 1 benchmark
WebQuestionsSP (WebQuestions Semantic Parses Dataset)
The WebQuestionsSP dataset is released as part of our ACL-2016 paper “The Value of Semantic Parse Labeling for Knowledge Base Question Answering” [Yih, Richardson, Meek, Chang & Suh, 2016], in which we evaluated the value of gathering…
61 papers · 3 benchmarks
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
ToTTo is an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description.
60 papers · 1 benchmark
CELEX database comprises three different searchable lexical databases, Dutch, English and German.
59 papers · 0 benchmarks
SParC (Semantic Parsing in Context)
SParC is a large-scale dataset for complex, cross-domain, and context-dependent (multi-turn) semantic parsing and text-to-SQL task (interactive natural language interfaces for relational databases).
59 papers · 2 benchmarks
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
CANARD (A Dataset for Question-in-Context Rewriting)
CANARD is a dataset for question-in-context rewriting that consists of questions each given in a dialog context together with a context-independent rewriting of the question.
58 papers · 1 benchmark
CCNet is a dataset extracted from Common Crawl with a different filtering process than for OSCAR.
58 papers · 0 benchmarks
EmoryNLP comprises 97 episodes, 897 scenes, and 12,606 utterances, where each utterance is annotated with one of the seven emotions borrowed from the six primary emotions in the Willcox (1982)’s feeling wheel, sad, mad, scared, powerful,…
58 papers · 1 benchmark
The Implicit Hate corpus is a dataset for hate speech detection with fine-grained labels for each message and its implication.
58 papers · 0 benchmarks
LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public.
58 papers · 2 benchmarks
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
SciDocs evaluation framework consists of a suite of evaluation tasks designed for document-level tasks.
57 papers · 2 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
MINC (Materials in Context Database)
MINC is a large-scale, open dataset of materials in the wild.
56 papers · 0 benchmarks
MasakhaNER is a collection of Named Entity Recognition (NER) datasets for 10 different African languages.
56 papers · 1 benchmark
RTE (Recognizing Textual Entailment)
The Recognizing Textual Entailment (RTE) datasets come from a series of textual entailment challenges.
56 papers · 2 benchmarks
Re-TACRED (Revised-TACRED)
The Re-TACRED dataset is a significantly improved version of the TACRED dataset for relation extraction.
56 papers · 1 benchmark
ISEAR (International Survey on Emotion Antecedents and Reactions)
Over a period of many years during the 1990s, a large group of psychologists all over the world collected data in the ISEAR project, directed by Klaus R.
55 papers · 0 benchmarks
MuTual is a retrieval-based dataset for multi-turn dialogue reasoning, which is modified from Chinese high school English listening comprehension test data.
55 papers · 0 benchmarks
PMC-VQA is a large-scale medical visual question-answering dataset that contains 227k VQA pairs of 149k images that cover various modalities or diseases.
55 papers · 2 benchmarks
PrOntoQA (Proof and Ontology-Generated Question-Answering)
PrOntoQA is a question-answering dataset which generates examples with chains-of-thought that describe the reasoning required to answer the questions correctly.
55 papers · 0 benchmarks
QUASAR-T (QUestion Answering by Search And Reading – Trivia)
QUASAR-T is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
55 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.