Home › Datasets › task › Text Generation

Text Generation datasets

archive 2025-07-28

162 datasets carry the task tag "Text Generation" (the task itself: Text Generation), ordered by the archive's paper count. Page 4 of 4: 18 shown of 162. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text Generation datasets 145–162 of 162

The RotoWire-Modified dataset is a cleaned extension of the RotoWire dataset, with writer information about each document.
1 paper · 0 benchmarks
SAD-Instruct (Situational Awareness Database for Instruct-Tuning)
The Situational Awareness Database for Instruct-Tuning (SAD-Instruct) is a dataset for dynamic task guidance.
1 paper · 0 benchmarks
SG-NLG (Schema-Guided Natural Language Generation)
The SG-NLG dataset is a pre-processed version of the DSTC8 Schema-Guided Dialogue SGD dataset, designed specifically for data-to-text Natural Language Generation (NLG).
1 paper · 0 benchmarks
SGXSTest (Singapore XSTest)
For testing refusal behavior in a cultural setting, we introduce SGXSTest — a set of manually curated prompts designed to measure exaggerated safety within the context of Singaporean culture.
1 paper · 0 benchmarks
In this Adjudicator ScoresShort Stories and Written Reflections folder: Four files from four student participants of the contest.
1 paper · 0 benchmarks
Dataset Card for "tamil-alpaca" This repository includes a Tamil-translated version of the Alpaca dataset.
1 paper · 0 benchmarks
Dataset Card for "tamil-alpaca" This repository includes a Tamil-translated versions of the Alpaca dataset and a subset of OpenOrca dataset.
1 paper · 0 benchmarks
TempWikiBio is a new data-to-text generation dataset containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles.
1 paper · 0 benchmarks
Texygen is a benchmarking platform to support research on open-domain text generation models.
1 paper · 0 benchmarks
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”.
1 paper · 0 benchmarks
UHGEvalDataset contains over 5000 news items.
1 paper · 0 benchmarks
Bangladesh's legal system struggles with major challenges like delays, complexity, high costs, and millions of unresolved cases, which deter many from pursuing legal action due to lack of knowledge or financial constraints.
1 paper · 0 benchmarks
WebBrain-Raw is a large-scale dataset built from English Wikipedia articles and their crawlable Wikipedia references.
1 paper · 0 benchmarks
WikiWeb2M (Wikipedia Webpage 2M)
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
We provide a new data set XWikiRef for the task of Cross-lingual Multi-document Summarization.
1 paper · 0 benchmarks
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
needadvice is a dataset for advice classification extracted from Reddit.
1 paper · 0 benchmarks
ASSIN2 (Avaliação de Similaridade Semântica e Inferência Textual) is the second edition of a workshop that evaluates Semantic Textual Similarity (STS) and Textual Entailment Recognition (RTE).
0 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.