Home › Datasets › task › Text Simplification

Text Simplification datasets

archive 2025-07-28

20 datasets carry the task tag "Text Simplification" (the task itself: Text Simplification), ordered by the archive's paper count. Page 1 of 1: 20 shown of 20. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text Simplification datasets 1–20 of 20

The Newsela dataset was introduced by Xu et al.
106 papers · 1 benchmark
WikiLarge comprise 359 test sentences, 2000 development sentences and 300k training sentences.
69 papers · 0 benchmarks
ASSET is a new dataset for assessing sentence simplification in English.
54 papers · 1 benchmark
TurkCorpus, a dataset with 2,359 original sentences from English Wikipedia, each with 8 manual reference simplifications.
44 papers · 1 benchmark
Useful for through two applications - automatic readability assessment and automatic text simplification.
29 papers · 0 benchmarks
Contains one million naturally occurring sentence rewrites, providing sixty times more distinct split examples and a ninety times larger vocabulary than the WebSplit corpus introduced by Narayan et al.
21 papers · 0 benchmarks
TextComplexityDE is a dataset consisting of 1000 sentences in German language taken from 23 Wikipedia articles in 3 different article-genres to be used for developing text-complexity predictor models and automatic text simplification in…
16 papers · 1 benchmark
CEFR-SP contains 17k English sentences annotated with the levels based on the Common European Framework of Reference for Languages assigned by English-education professionals.
7 papers · 0 benchmarks
Klexikon (Klexikon: A German Dataset for Joint Summarization and Simplification)
The dataset introduces document alignments between German Wikipedia and the children's lexicon Klexikon.
5 papers · 1 benchmark
EurekaAlert (Eureka Alert)
This dataset contains around 5000 scholarly articles and their corresponding easy summary from eureka alert blog, the dataset can be used for the combined task of summarization and simplification.
4 papers · 2 benchmarks
Med-EASi (Medical dataset for Elaborative and Abstractive Simplification), a uniquely crowdsourced and finely annotated dataset for supervised simplification of short medical texts.
4 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
DEplain-APA-sent: A German Parallel Corpus for Sentence Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
DEplain-web-sent: A German Parallel Corpus for Sentence Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
DEplain-APA-doc: A German Parallel Corpus for Document Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
DEplain-web-doc: A German Parallel Corpus for Document Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
1 paper · 0 benchmarks
Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature.
1 paper · 0 benchmarks
A medical Wiki paralell corpus for medical text simplification.
1 paper · 0 benchmarks
SimpEvalASSET is a dataset for learning learnable metrics using modern language models.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.