Home › Datasets › task › Text Generation
Text Generation datasets
archive 2025-07-28
162 datasets carry the task tag "Text Generation" (the task itself: Text Generation), ordered by the archive's paper count. Page 2 of 4: 48 shown of 162. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Text Generation datasets 49–96 of 162
KELM is a large-scale synthetic corpus of Wikidata KG as natural text.
24 papers · 0 benchmarks
Moral Stories is a crowd-sourced dataset of structured narratives that describe normative and norm-divergent actions taken by individuals to accomplish certain intentions in concrete situations, and their respective consequences.
24 papers · 0 benchmarks
HeadQA is a multi-choice question answering testbed to encourage research on complex reasoning.
23 papers · 1 benchmark
AMR Bank (Abstract Meaning Representation)
The AMR Bank is a set of English sentences paired with simple, readable semantic representations.
22 papers · 1 benchmark
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
WIQA (What-If Question Answering)
The WIQA dataset V1 has 39705 questions containing a perturbation and a possible effect in the context of a paragraph.
21 papers · 0 benchmarks
RoboCup is an initiative in which research groups compete by enabling their robots to play football matches.
20 papers · 0 benchmarks
IndicGLUE (Indic General Language Understanding Evaluation Benchmark)
We now introduce IndicGLUE, the Indic General Language Understanding Evaluation Benchmark, which is a collection of various NLP tasks as de- scribed below.
16 papers · 4 benchmarks
CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy.
15 papers · 0 benchmarks
Opusparcus is a paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish.
15 papers · 0 benchmarks
AmazonQA consists of 923k questions, 3.6M answers and 14M reviews across 156k products.
14 papers · 0 benchmarks
TimeTravel contains 29,849 counterfactual rewritings, each with the original story, a counterfactual event, and human-generated revision of the original story compatible with the counterfactual event.
14 papers · 1 benchmark
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
ROSCOE is a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics.
13 papers · 0 benchmarks
OPUS (open parallel corpus)
OPUS is a growing collection of translated texts from the web.
12 papers · 0 benchmarks
CC-Stories (or STORIES) is a dataset for common sense reasoning and language modeling.
10 papers · 0 benchmarks
Chart2Text is a dataset that was crawled from 23,382 freely accessible pages from statista.com in early March of 2020, yielding a total of 8,305 charts, and associated summaries.
10 papers · 0 benchmarks
ChatHaruhi (ChatHaruhi: Reviving Anime Character in Reality via Large Language Model)
ChatHaruhi is a dataset covering 32 Chinese / English TV / anime characters with over 54k simulated dialogues.
10 papers · 0 benchmarks
WikiAtomicEdits is a corpus of 43 million atomic edits across 8 languages.
10 papers · 0 benchmarks
Paper | Github | Dataset| Model As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e.
9 papers · 1 benchmark
Logic2Text is a large-scale dataset with 10,753 descriptions involving common logic types paired with the underlying logical forms.
8 papers · 0 benchmarks
LongForm dataset is created by leveraging English corpus examples with augmented instructions.
8 papers · 0 benchmarks
WikiTableT contains Wikipedia article sections and their corresponding tabular data and various metadata.
8 papers · 0 benchmarks
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
CELLS is a large (63k pairs) and broadest-ranging (12 journals) parallel corpus for lay language generation.
6 papers · 0 benchmarks
CLUECorpus2020 is a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation.
6 papers · 0 benchmarks
Microsoft Research Social Media Conversation Corpus consists of 127M context-message-response triples from the Twitter FireHose, covering the 3-month period June 2012 through August 2012.
6 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
Benchmark dataset for abstracts and titles of 100,000 ArXiv scientific papers.
6 papers · 1 benchmark
A new dataset on the baseball domain.
5 papers · 1 benchmark
PoMo consists of more than 231K sentences with post-modifiers and associated facts extracted from Wikidata for around 57K unique entities.
5 papers · 0 benchmarks
RiSAWOZ is a large-scale multi-domain Chinese Wizard-of-Oz dataset with Rich Semantic Annotations.
5 papers · 0 benchmarks
The Russian Corpus of Linguistic Acceptability (RuCoLA) is built from the ground up under the well-established binary LA approach.
5 papers · 1 benchmark
STACKEX expands beyond the only existing genre (i.e., academic writing) in keyphrase generation tasks.
5 papers · 0 benchmarks
A multilingual image dataset with spatial relation annotations and object features for image-to-text generation, built using 2,026 images from the PASCAL VOC2008 dataset.
5 papers · 0 benchmarks
Taiga is a corpus, where text sources and their meta-information are collected according to popular ML tasks.
5 papers · 0 benchmarks
The TweetSentBR Dataset is a valuable resource for sentiment analysis in Brazilian Portuguese.
5 papers · 1 benchmark
A new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue.
4 papers · 1 benchmark
Goal is a novel dataset of football (or 'soccer') highlights videos with transcribed live commentaries in English.
4 papers · 0 benchmarks
Wikipedia Generation is a dataset for article generation from Wikipedia from references at the end of Wikipedia page and the top 10 search results for the Wikipedia topic.
4 papers · 0 benchmarks
CANNOT (Compilation of ANnotated, Negation-Oriented Text-pairs)
Dataset Summary CANNOT is a dataset that focuses on negated textual pairs.
3 papers · 0 benchmarks
DR.BENCH (Diagnostic Reasoning Benchmark for clinical natural language processing)
DR.BENCH is a dataset for developing and evaluating cNLP models with clinical diagnostic reasoning ability.
3 papers · 0 benchmarks
The OAB Exams dataset is a valuable resource used in the context of legal information systems.
3 papers · 1 benchmark
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
Hugging Face Datasets (New!) | Website | Github Repository | arXiv e-Print The Visual Writing Prompts (VWP) dataset contains almost 2K selected sequences of movie shots, each including 5-10 images.
3 papers · 0 benchmarks
A dataset of single-sentence edits crawled from Wikipedia.
3 papers · 0 benchmarks
YTD-18M is a large-scale corpus of 18M video-based dialogues, constructed from web videos: crucial to the data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format…
3 papers · 0 benchmarks
CLSE (Corpus of Linguistically Significant Entities)
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.