Home › Datasets › task › Abstractive Text Summarization

Abstractive Text Summarization datasets

archive 2025-07-28

52 datasets carry the task tag "Abstractive Text Summarization" (the task itself: Abstractive Text Summarization), ordered by the archive's paper count. Page 1 of 2: 48 shown of 52. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Abstractive Text Summarization datasets 1–48 of 52

The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
WikiHow is a dataset of more than 230,000 article and summary pairs extracted and constructed from an online knowledge base written by different human authors.
127 papers · 2 benchmarks
NEWSROOM (CORNELL NEWSROOM)
CORNELL NEWSROOM is a large dataset for training and evaluating summarization systems.
107 papers · 0 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.
62 papers · 2 benchmarks
LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public.
58 papers · 2 benchmarks
WikiSum is a dataset based on English Wikipedia and suitable for a task of multi-document abstractive summarization.
54 papers · 0 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
BillSum is the first dataset for summarization of US Congressional and California state bills.
47 papers · 2 benchmarks
Reddit TIFU dataset is a newly collected Reddit dataset, where TIFU denotes the name of /r/tifu subbreddit.
47 papers · 1 benchmark
MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
BookSum is a collection of datasets for long-form narrative summarization.
39 papers · 2 benchmarks
The Extreme Summarization (XSum) dataset is a dataset for evaluation of abstractive single-document summarization systems.
32 papers · 5 benchmarks
eLife (Scientific Lay Summarization)
This dataset contains 4,828 full biomedical articles paired with non-technical lay summaries derived from the eLife scientific journal.
24 papers · 2 benchmarks
To study the task of email subject line generation: automatically generating an email subject line from the email body.
22 papers · 1 benchmark
AMR Bank (Abstract Meaning Representation)
The AMR Bank is a set of English sentences paired with simple, readable semantic representations.
22 papers · 1 benchmark
PLOS (Scientific Lay Summarization)
This dataset contains 27,525 full biomedical articles paired with non-technical lay summaries derived from various journals published by the Public Library of Science (PLOS).
21 papers · 2 benchmarks
A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
13 papers · 0 benchmarks
This is a dataset for evaluating summarisation methods for research papers.
13 papers · 3 benchmarks
GLGE (General Language Generation Evaluation)
GLGE is a general language generation evaluation benchmark which is composed of 8 language generation tasks, including Abstractive Text Summarization (CNN/DailyMail, Gigaword, XSUM, MSNews), Answer-aware Question Generation (SQuAD 1.1,…
12 papers · 0 benchmarks
5,519 query-based summaries, each associated with an average of 6 input documents selected from an index of 355M documents from Common Crawl.
11 papers · 0 benchmarks
CRD3 (Critical Role Dungeons and Dragons Dataset)
The dataset is collected from 159 Critical Role episodes transcribed to text dialogues, consisting of 398,682 turns.
9 papers · 0 benchmarks
Global Voices is a multilingual dataset for evaluating cross-lingual summarization methods.
9 papers · 0 benchmarks
NCLS (Neural Cross-Lingual Summarization Corpora)
Presents two high-quality large-scale CLS datasets based on existing monolingual summarization datasets.
9 papers · 0 benchmarks
ConvoSumm is a suite of four datasets to evaluate a model’s performance on a broad spectrum of conversation data.
6 papers · 0 benchmarks
A set of approximately 100K podcast episodes comprised of raw audio files along with accompanying ASR transcripts.
6 papers · 0 benchmarks
A large-scale Indonesian summarization dataset consisting of harvested articles from Liputan6.com, an online news portal, resulting in 215,827 document-summary pairs.
5 papers · 0 benchmarks
Pn-summary is a dataset for Persian abstractive text summarization.
4 papers · 0 benchmarks
DMQA (DeepMind Q&A)
The DeepMind Q&A Dataset consists of two datasets for Question Answering, CNN and DailyMail.
3 papers · 0 benchmarks
NarraSum is a large-scale narrative summarization dataset.
3 papers · 0 benchmarks
Shmoop Corpus is a dataset of 231 stories that are paired with detailed multi-paragraph summaries for each individual chapter (7,234 chapters), where the summary is chronologically aligned with respect to the story chapter.
3 papers · 0 benchmarks
Fanpage dataset, containing news articles taken from Fanpage.
2 papers · 1 benchmark
IlPost dataset, containing news articles taken from IlPost.
2 papers · 1 benchmark
FINDSum (Financial Report Document Summarization)
FINDSum is a large-scale dataset for long text and multi-table summarization.
2 papers · 0 benchmarks
PeerSum is a new MDS dataset using peer reviews of scientific publications.
2 papers · 0 benchmarks
PoC (Points of correspondence)
A dataset containing the documents, source and fusion sentences, and human annotations of points of correspondence between sentences.
2 papers · 0 benchmarks
VNDS (VNDS: A Vietnamese Dataset for Summarization)
A single-document Vietnamese summarization dataset
2 papers · 1 benchmark
This is a dataset for multi-document summarization in Portuguese, what means that it has examples of multiple documents (input) related to human-written summaries (output).
1 paper · 0 benchmarks
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline.
1 paper · 0 benchmarks
ICLR Database (ICLR Database (with Textual Covariates))
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
1 paper · 0 benchmarks
Inshorts News (Inshorts English News dataset)
Inshorts News dataset Inshorts provides a news summary in 60 words or less.
1 paper · 1 benchmark
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
1 paper · 1 benchmark
MOS Dataset (Microblog Opinion Summarisation)
This dataset was used in the paper 'Template-based Abstractive Microblog Opinion Summarisation' (to be published at TACL, 2022).
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.