Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 1 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1–48 of 3,130
GLUE (General Language Understanding Evaluation benchmark)
General Language Understanding Evaluation (GLUE) benchmark is a collection of nine natural language understanding tasks, including single-sentence tasks CoLA and SST-2, similarity and paraphrasing tasks MRPC, STS-B and QQP, and natural…
3,197 papers · 13 benchmarks
SST (Stanford Sentiment Treebank)
The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language.
2,354 papers · 6 benchmarks
MML (Massive Multitask Language Understanding)
MMLU (Massive Multitask Language Understanding) is a new benchmark designed to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings.
1,922 papers · 29 benchmarks
GSM8K is a dataset of 8.5K high quality linguistically diverse grade school math word problems created by human problem writers.
1,881 papers · 7 benchmarks
Visual Question Answering (VQA) is a dataset containing open-ended questions about images.
1,834 papers · 0 benchmarks
MultiNLI (Multi-Genre Natural Language Inference)
The Multi-Genre Natural Language Inference (MultiNLI) dataset has 433K sentence pairs.
1,830 papers · 4 benchmarks
The IMDb Movie Reviews dataset is a binary sentiment analysis dataset consisting of 50,000 reviews from the Internet Movie Database (IMDb) labeled as positive or negative.
1,787 papers · 9 benchmarks
The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples.
1,404 papers · 9 benchmarks
MATH is a new dataset of 12,500 challenging competition mathematics problems.
1,330 papers · 2 benchmarks
SNLI (Stanford Natural Language Inference)
The SNLI dataset (Stanford Natural Language Inference) consists of 570k sentence-pairs manually labeled as entailment, contradiction, and neutral.
1,311 papers · 1 benchmark
Visual Genome contains Visual Question Answering data in a multi-choice setting.
1,256 papers · 15 benchmarks
QNLI (Question-answering NLI)
The QNLI (Question-answering NLI) dataset is a Natural Language Inference dataset automatically derived from the Stanford Question Answering Dataset v1.1 (SQuAD).
1,234 papers · 3 benchmarks
This is an evaluation harness for the HumanEval problem solving dataset described in the paper "Evaluating Large Language Models Trained on Code".
1,201 papers · 1 benchmark
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
1,081 papers · 1 benchmark
MS MARCO (Microsoft Machine Reading Comprehension Dataset)
The MS MARCO (Microsoft MAchine Reading Comprehension) is a collection of datasets focused on deep learning in search.
1,036 papers · 7 benchmarks
The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling.
1,006 papers · 10 benchmarks
C4 (Colossal Clean Crawled Corpus)
C4 is a colossal, cleaned version of Common Crawl's web crawl corpus.
981 papers · 1 benchmark
AG News (AG’s News Corpus) is a subdataset of AG's corpus of news articles constructed by assembling titles and description fields of articles from the 4 largest classes (“World”, “Sports”, “Business”, “Sci/Tech”) of AG’s Corpus.
969 papers · 9 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
HotpotQA is a question answering dataset collected on the English Wikipedia, containing about 113K crowd-sourced questions that are constructed to require the introduction paragraphs of two Wikipedia articles to answer.
933 papers · 3 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
ConceptNet is a knowledge graph that connects words and phrases of natural language with labeled edges.
852 papers · 1 benchmark
MRPC (Microsoft Research Paraphrase Corpus)
Microsoft Research Paraphrase Corpus (MRPC) is a corpus consists of 5,801 sentence pairs collected from newswire articles.
786 papers · 4 benchmarks
PIQA (Physical Interaction: Question Answering)
PIQA is a dataset for commonsense reasoning, and was created to investigate the physical knowledge of existing models in NLP.
772 papers · 2 benchmarks
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs.
749 papers · 8 benchmarks
CoLA (Corpus of Linguistic Acceptability)
The Corpus of Linguistic Acceptability (CoLA) consists of 10657 sentences from 23 linguistics publications, expertly annotated for acceptability (grammaticality) by their original authors.
710 papers · 4 benchmarks
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples.
701 papers · 5 benchmarks
CLEVR (Compositional Language and Elementary Visual Reasoning)
CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset.
657 papers · 3 benchmarks
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject.
635 papers · 3 benchmarks
DBpedia (from "DB" for "database") is a project aiming to extract structured content from the information created in the Wikipedia project.
597 papers · 4 benchmarks
VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media.
564 papers · 5 benchmarks
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
560 papers · 2 benchmarks
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
531 papers · 1 benchmark
CNN/Daily Mail is a dataset for text summarization.
530 papers · 8 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
FEVER (Fact Extraction and VERification)
FEVER is a publicly available dataset for fact extraction and verification against textual sources.
498 papers · 3 benchmarks
The CommonsenseQA is a dataset for commonsense question answering task.
483 papers · 1 benchmark
TextVQA is a dataset to benchmark visual reasoning based on text in images.
476 papers · 3 benchmarks
This CSTR VCTK Corpus includes speech data uttered by 110 English speakers with various accents.
476 papers · 6 benchmarks
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.
467 papers · 1 benchmark
The RefCOCO dataset is a referring expression generation (REG) dataset used for tasks related to understanding natural language expressions that refer to specific objects in images.
439 papers · 11 benchmarks
WebText is an internal OpenAI corpus created by scraping web pages with emphasis on document quality.
425 papers · 0 benchmarks
SuperGLUE is a benchmark dataset designed to pose a more rigorous test of language understanding than GLUE.
423 papers · 0 benchmarks
RACE (ReAding Comprehension dataset from Examinations)
The ReAding Comprehension dataset from Examinations (RACE) dataset is a machine reading comprehension dataset consisting of 27,933 passages and 97,867 questions from English exams, targeting Chinese students aged 12-18.
412 papers · 3 benchmarks
DailyDialog is a high-quality multi-turn open-domain English dialog dataset.
399 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.