Home › Datasets › task › Text Generation
Text Generation datasets
archive 2025-07-28
162 datasets carry the task tag "Text Generation" (the task itself: Text Generation), ordered by the archive's paper count. Page 1 of 4: 48 shown of 162. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Text Generation datasets 1–48 of 162
GLUE (General Language Understanding Evaluation benchmark)
General Language Understanding Evaluation (GLUE) benchmark is a collection of nine natural language understanding tasks, including single-sentence tasks CoLA and SST-2, similarity and paraphrasing tasks MRPC, STS-B and QQP, and natural…
3,197 papers · 13 benchmarks
MML (Massive Multitask Language Understanding)
MMLU (Massive Multitask Language Understanding) is a new benchmark designed to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings.
1,922 papers · 29 benchmarks
GSM8K is a dataset of 8.5K high quality linguistically diverse grade school math word problems created by human problem writers.
1,881 papers · 7 benchmarks
MultiNLI (Multi-Genre Natural Language Inference)
The Multi-Genre Natural Language Inference (MultiNLI) dataset has 433K sentence pairs.
1,830 papers · 4 benchmarks
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
1,081 papers · 1 benchmark
HellaSwag is a challenge dataset for evaluating commonsense NLI that is specially hard for state-of-the-art models, though its questions are trivial for humans (>95% accuracy).
994 papers · 6 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
PIQA (Physical Interaction: Question Answering)
PIQA is a dataset for commonsense reasoning, and was created to investigate the physical knowledge of existing models in NLP.
772 papers · 2 benchmarks
CoLA (Corpus of Linguistic Acceptability)
The Corpus of Linguistic Acceptability (CoLA) consists of 10657 sentences from 23 linguistics publications, expertly annotated for acceptability (grammaticality) by their original authors.
710 papers · 4 benchmarks
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples.
701 papers · 5 benchmarks
OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject.
635 papers · 3 benchmarks
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions.
607 papers · 3 benchmarks
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
560 papers · 2 benchmarks
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
531 papers · 1 benchmark
CNN/Daily Mail is a dataset for text summarization.
530 papers · 8 benchmarks
DailyDialog is a high-quality multi-turn open-domain English dialog dataset.
399 papers · 2 benchmarks
DROP (Discrete Reasoning Over Paragraphs)
Discrete Reasoning Over Paragraphs DROP is a crowdsourced, adversarially-created, 96k-question benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over…
382 papers · 3 benchmarks
BookCorpus is a large collection of free novel books written by unpublished authors, which contains 11,038 books (around 74M sentences and 1G words) of 16 different sub-genres (e.g., Romance, Historical, Adventure, etc.).
344 papers · 1 benchmark
OpenWebText is an open-source recreation of the WebText corpus.
207 papers · 2 benchmarks
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
COCO Captions contains over one and a half million captions describing over 330,000 images.
203 papers · 4 benchmarks
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others.
183 papers · 1 benchmark
ELI5 is a dataset for long-form question answering.
158 papers · 1 benchmark
ROCStories is a collection of commonsense short stories.
152 papers · 2 benchmarks
The One Billion Word dataset is a dataset for language modeling.
141 papers · 2 benchmarks
The AlpacaEval set contains 805 instructions form self-instruct, open-assistant, vicuna, koala, hh-rlhf.
131 papers · 2 benchmarks
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
WritingPrompts is a large dataset of 300K human-written stories paired with writing prompts from an online forum.
118 papers · 1 benchmark
The Stack contains over 3TB of permissively-licensed source code files covering 30 programming languages crawled from GitHub.
115 papers · 0 benchmarks
CommonGen is constructed through a combination of crowdsourced and existing caption corpora, consists of 79k commonsense descriptions over 35k unique concept-sets.
108 papers · 1 benchmark
ReDial (Recommendation Dialogues) is an annotated dataset of dialogues, where users recommend movies to each other.
105 papers · 2 benchmarks
E2E (End-to-End NLG Challenge)
End-to-End NLG Challenge (E2E) aims to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena.
90 papers · 4 benchmarks
FLoRes-101 is an evaluation benchmark for low-resource and multilingual machine translation.
86 papers · 57 benchmarks
RealNews is a large corpus of news articles from Common Crawl.
80 papers · 0 benchmarks
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
Sentence Compression is a dataset where the syntactic trees of the compressions are subtrees of their uncompressed counterparts, and hence where supervised systems which require a structural alignment between the input and output can be…
63 papers · 0 benchmarks
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
62 papers · 0 benchmarks
LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public.
58 papers · 2 benchmarks
MuTual is a retrieval-based dataset for multi-turn dialogue reasoning, which is modified from Chinese high school English listening comprehension test data.
55 papers · 0 benchmarks
OpenDialKG contains utterance from 15K human-to-human role-playing dialogs is manually annotated with ground-truth reference to corresponding entities and paths from a large-scale KG with 1M+ facts.
55 papers · 0 benchmarks
The Machine Translation of Noisy Text (MTNT) dataset is a Machine Translation dataset that consists of noisy comments on Reddit and professionally sourced translation.
52 papers · 0 benchmarks
DART is a large dataset for open-domain structured data record to text generation.
45 papers · 3 benchmarks
CCMatrix uses ten snapshots of a curated common crawl corpus (Wenzek et al., 2019) totalling 32.7 billion unique sentences.
42 papers · 0 benchmarks
CSL is a synthetic dataset introduced in Murphy et al.
34 papers · 2 benchmarks
The Image Paragraph Captioning dataset allows researchers to benchmark their progress in generating paragraphs that tell a story about an image.
31 papers · 1 benchmark
CONAN (COunter NArratives through Nichesourcing)
COunter NArratives through Nichesourcing (CONAN) is a dataset that consists of 4,078 pairs over the 3 languages.
27 papers · 0 benchmarks
KPTimes is a large-scale dataset of news texts paired with editor-curated keyphrases.
27 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.