Home › Datasets › task › Text Generation

Text Generation datasets

archive 2025-07-28

162 datasets carry the task tag "Text Generation" (the task itself: Text Generation), ordered by the archive's paper count. Page 3 of 4: 48 shown of 162. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text Generation datasets 97–144 of 162

CodeSyntax is a large-scale dataset of programs annotated with the syntactic relationships in their corresponding abstract syntax trees.
2 papers · 0 benchmarks
Concise has two datasets of 2000 sentences each, that were annotated by two and five human annotators, respectively.
2 papers · 0 benchmarks
Czech restaurant information is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language.
2 papers · 1 benchmark
DIALOCONAN is a dataset comprising over 3000 fictitious multi-turn dialogues between a hater and an NGO operator, covering 6 targets of hate.
2 papers · 0 benchmarks
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
ExHVV is a novel dataset that offers natural language explanations of connotative roles for three types of entities -- heroes, villains, and victims, encompassing 4,680 entities present in 3K memes.
2 papers · 0 benchmarks
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
The Live Comment Dataset is a large-scale dataset with 2,361 videos and 895,929 live comments that were written while the videos were streamed.
2 papers · 0 benchmarks
This dataset consists of 5808 dialogues, based on 2236 unique scenarios.
2 papers · 0 benchmarks
The QTUNA dataset is the result of a series of elicitation experiments in which human speakers were asked to perform a linguistic task that invites the use of quantified expressions in order to inform possible Natural Language Generation…
2 papers · 0 benchmarks
The TaoDescribe dataset contains 2,129,187 product titles and descriptions in Chinese.
2 papers · 0 benchmarks
ThreatGram 101 - Extreme Telegram Data (ThreatGram 101 - Extreme Telegram Replies Data with Threat Levels)
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
AAVE/SAE Paired Dataset contains 2019 intent-equivalent AAVE/SAE pairs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The AppealCase dataset is the first large-scale resource specifically designed to support LegalAI research in appellate judgment scenarios.
1 paper · 0 benchmarks
AskParents is a dataset for advice classification extracted from Reddit.
1 paper · 0 benchmarks
A syllogism is a common form of deductive reasoning that requires precisely two premises and one conclusion.
1 paper · 0 benchmarks
COLLIE-v1 is a dataset with 2080 instances comprising 13 constraint structures designed for text generation under constraints.
1 paper · 0 benchmarks
A large dataset of color names and their respective RGB values stores in CSV.
1 paper · 1 benchmark
DocBank-TB (DocBank-Table)
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
1 paper · 0 benchmarks
DpgMedia2019 is a Dutch news dataset for partisanship detection.
1 paper · 0 benchmarks
ENT-DESC involves retrieving abundant knowledge of various types of main entities from a large knowledge graph (KG), which makes the current graph-to-sequence models severely suffer from the problems of information loss and parameter…
1 paper · 1 benchmark
Educational Grade School Math (EGSM) contains 2,093 question/answer pairs generated by MATHWELL, a reference-free educational grade school math word problem generator that outputs a word problem and Program of Thought (PoT) solution based…
1 paper · 0 benchmarks
Expository Prose (Expository-Prose-V1)
Expository-Prose-V1 is a collection of specially-curated corpora gathered from diverse sources, ranging from research papers (arXiv) to European Parliament proceedings (EuroParl).
1 paper · 0 benchmarks
Food.com Recipes and Interactions consists of 270K recipes and 1.4M user-recipe interactions (reviews) scraped from Food.com, covering a period of 18 years (January 2000 to December 2018).
1 paper · 0 benchmarks
Fraud_Case_Verdicts (The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset)
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HAVOC (Harmful Abstractions and Violations in Open Completions Benchmark)
measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes.
1 paper · 0 benchmarks
HiXSTest (Hindi XSTest)
For testing refusal behavior in a language-specific setting, we introduce HiXSTest — a set of manually curated prompts in the Hindi language designed to measure exaggerated safety.
1 paper · 0 benchmarks
A collection of test sets for evaluating base and chat LLMs (incl.
1 paper · 0 benchmarks
Ice Hockey News Dataset is a corpus of Finnish ice hockey news, edited to be suitable for training of end-to-end news generation methods, as well as demonstrate generation of text, which was judged by journalists to be relatively close to…
1 paper · 0 benchmarks
Image Caption Quality Dataset is a dataset of crowdsourced ratings for machine-generated image captions.
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The Lenta Short Sentences dataset is a text dataset for language modelling for the Russian language.
1 paper · 0 benchmarks
This is a dataset of 3 English books which do not contain the letter "e" in them.
1 paper · 1 benchmark
An open-source online generative dictionary that takes a word and context containing the word as input and automatically generates a definition as output.
1 paper · 0 benchmarks
MTTN is a large scale derived and synthesized dataset built with on real prompts and indexed with popular image-text datasets like MS-COCO, Flickr, etc.
1 paper · 0 benchmarks
Dataset introduction There are four dimension in MBTI.
1 paper · 0 benchmarks
pymatgencodeqa benchmark: qabenchmark/generatedqa/generationresultscode.json, which consists of 34,621 QA pairs.
1 paper · 0 benchmarks
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
The OnlySports Dataset is a comprehensive collection of sports-related text data, comprising approximately 600 billion tokens.
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
Pentachromatic Cultural Palette Dataset is characterized by unique cultural semantics and values.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
Dataset Details Total Labeled: 100% Labeled and Curated: 24,478 Pending: 0 Drafts: 0 Discarded: 696 High-Level Explanation This dataset includes labeled samples from the Colombian Aeronautical Regulations (RAC), covering all chapters…
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.