Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 4 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 145–192 of 3,130

A-OKVQA is crowdsourced visual question answering dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer.
154 papers · 1 benchmark
The NCBI Disease corpus consists of 793 PubMed abstracts, which are separated into training (593), development (100) and test (100) subsets.
154 papers · 3 benchmarks
CC12M (Conceptual 12M)
Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training.
153 papers · 0 benchmarks
OLID (Offensive Language Identification Dataset)
The OLID is a hierarchical dataset to identify the type and the target of offensive texts in social media.
152 papers · 1 benchmark
ROCStories is a collection of commonsense short stories.
152 papers · 2 benchmarks
A large corpus of 81.1M English-language academic papers spanning many academic disciplines.
152 papers · 1 benchmark
Video-MME stands for Video Multi-Modal Evaluation.
152 papers · 2 benchmarks
FCE (First Certificate in English)
The Cambridge Learner Corpus First Certificate in English (CLC FCE) dataset consists of short texts, written by learners of English as an additional language in response to exam prompts eliciting free-text answers and assessing mastery of…
151 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
VRD (Visual Relationship Detection dataset)
The Visual Relationship Dataset (VRD) contains 4000 images for training and 1000 for testing annotated with visual relationships.
151 papers · 5 benchmarks
The WebNLG corpus comprises of sets of triplets describing facts (entities and relations between them) and the corresponding facts in form of natural language text.
149 papers · 17 benchmarks
SCAN (Simplified versions of the CommAI Navigation tasks)
SCAN is a dataset for grounded navigation which consists of a set of simple compositional navigation commands paired with the corresponding action sequences.
148 papers · 0 benchmarks
TyDiQA (Typologically Diverse Question Answering)
TyDi QA is a question answering dataset covering 11 typologically diverse languages with 200K question-answer pairs.
148 papers · 0 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
The TVQA dataset is a large-scale video dataset for video question answering.
146 papers · 3 benchmarks
VQA-RAD (Visual Question Answering in Radiology)
VQA-RAD consists of 3,515 question–answer pairs on 315 radiology images.
145 papers · 0 benchmarks
Wizard of Wikipedia is a large dataset with conversations directly grounded with knowledge retrieved from Wikipedia.
145 papers · 1 benchmark
MedMCQA is a large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions.
144 papers · 1 benchmark
The Flickr30K Entities dataset is an extension to the Flickr30K dataset.
142 papers · 2 benchmarks
mC4 is a multilingual variant of the C4 dataset called mC4.
142 papers · 0 benchmarks
The One Billion Word dataset is a dataset for language modeling.
141 papers · 2 benchmarks
FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech)
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark.
141 papers · 1 benchmark
e-SNLI is used for various goals, such as obtaining full sentence justifications of a model's decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets.
139 papers · 1 benchmark
SciERC dataset is a collection of 500 scientific abstract annotated with scientific entities, their relations, and coreference clusters.
134 papers · 7 benchmarks
WinoBias contains 3,160 sentences, split equally for development and test, created by researchers familiar with the project.
134 papers · 0 benchmarks
BLUE (Biomedical Language Understanding Evaluation)
The BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
133 papers · 0 benchmarks
SearchQA was built using an in-production, commercial search engine.
133 papers · 1 benchmark
The AlpacaEval set contains 805 instructions form self-instruct, open-assistant, vicuna, koala, hh-rlhf.
131 papers · 2 benchmarks
GoEmotions is a corpus of 58k carefully curated comments extracted from Reddit, with human annotations to 27 emotion categories or Neutral.
130 papers · 0 benchmarks
LIAR is a publicly available dataset for fake news detection.
130 papers · 1 benchmark
TabFact is a large-scale dataset which consists of 117,854 manually annotated statements with regard to 16,573 Wikipedia tables, their relations are classified as ENTAILED and REFUTED.
129 papers · 2 benchmarks
Europarl (European Parliament Proceedings Parallel Corpus)
A corpus of parallel text in 21 European languages from the proceedings of the European Parliament.
128 papers · 1 benchmark
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
WNUT 2017 (WNUT 2017 Emerging and Rare entity recognition)
This shared task focuses on identifying unusual, previously-unseen entities in the context of emerging discussions.
127 papers · 2 benchmarks
WikiHow is a dataset of more than 230,000 article and summary pairs extracted and constructed from an online knowledge base written by different human authors.
127 papers · 2 benchmarks
BLiMP (Benchmark of Linguistic Minimal Pairs)
BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English.
126 papers · 0 benchmarks
LSMDC (Large Scale Movie Description Challenge)
This dataset contains 118,081 short video clips extracted from 202 movies.
126 papers · 3 benchmarks
MSRA-TD500 (MSRA Text Detection 500 Database)
The MSRA-TD500 dataset is a text detection dataset that contains 300 training images and 200 test images.
124 papers · 1 benchmark
The Microsoft Academic Graph is a heterogeneous graph containing scientific publication records, citation relationships between those publications, as well as authors, institutions, journals, conferences, and fields of study.
124 papers · 0 benchmarks
CARER (Contextualized Affect Representations for Emotion Recognition)
CARER is an emotion dataset collected through noisy labels, annotated via distant supervision as in (Go et al., 2009).
123 papers · 2 benchmarks
BBQ (Bias Benchmark for QA)
Bias Benchmark for QA (BBQ) is a dataset consisting of question-sets constructed by the authors that highlight attested social biases against people belonging to protected classes along nine different social dimensions relevant for U.S.
122 papers · 0 benchmarks
Multi-News, consists of news articles and human-written summaries of these articles from the site newser.com.
122 papers · 5 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
SIQA (Social Interaction QA)
Social Interaction QA (SIQA) is a question-answering benchmark for testing social commonsense intelligence.
120 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.