Home › Datasets › task › Document Summarization
Document Summarization datasets
archive 2025-07-28
28 datasets carry the task tag "Document Summarization" (the task itself: Document Summarization), ordered by the archive's paper count. Page 1 of 1: 28 shown of 28. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Document Summarization datasets 1–28 of 28
CNN/Daily Mail is a dataset for text summarization.
530 papers · 8 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
Multi-News, consists of news articles and human-written summaries of these articles from the site newser.com.
122 papers · 5 benchmarks
CORNELL NEWSROOM is a large dataset for training and evaluating summarization systems.
107 papers · 0 benchmarks
WikiSum is a dataset based on English Wikipedia and suitable for a task of multi-document abstractive summarization.
54 papers · 0 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
BookSum is a collection of datasets for long-form narrative summarization.
39 papers · 2 benchmarks
Arxiv HEP-TH (high energy physics theory) citation graph is from the e-print arXiv and covers all the citations within a dataset of 27,770 papers with 352,807 edges.
35 papers · 5 benchmarks
WCEP (Wikipedia Current Events Portal)
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an…
31 papers · 1 benchmark
To study the task of email subject line generation: automatically generating an email subject line from the email body.
22 papers · 1 benchmark
Large-scale manually-annotated corpus for 1,000 scientific papers (on computational linguistics) for automatic summarization.
20 papers · 0 benchmarks
This is a dataset for evaluating summarisation methods for research papers.
13 papers · 3 benchmarks
5,519 query-based summaries, each associated with an average of 6 input documents selected from an index of 355M documents from Common Crawl.
11 papers · 0 benchmarks
Multi-XScience is a large-scale dataset for multi-document summarization of scientific articles.
10 papers · 0 benchmarks
CQASUMM is a dataset for CQA (Community Question Answering) summarization, constructed from the 4.4 million Yahoo!
9 papers · 0 benchmarks
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks
OPOSUM is a dataset for the training and evaluation of Opinion Summarization models which contains Amazon reviews from six product domains: Laptop Bags, Bluetooth Headsets, Boots, Keyboards, Televisions, and Vacuums.
7 papers · 0 benchmarks
The TalkSumm dataset contains 1705 automatically-generated summaries of scientific papers from ACL, NAACL, EMNLP, SIGDIAL (2015-2018), and ICML (2017-2018).
6 papers · 0 benchmarks
MATINF (Maternal and Infant Dataset)
Maternal and Infant (MATINF) Dataset is a large-scale dataset jointly labeled for classification, question answering and summarization in the domain of maternity and baby caring in Chinese.
5 papers · 0 benchmarks
FacetSum is a faceted summarization dataset for scientific documents.
4 papers · 1 benchmark
Wikipedia Generation is a dataset for article generation from Wikipedia from references at the end of Wikipedia page and the top 10 search results for the Wikipedia topic.
4 papers · 0 benchmarks
An open corpus of Scientific Research papers which has a representative sample from across scientific disciplines.
2 papers · 0 benchmarks
FINDSum (Financial Report Document Summarization)
FINDSum is a large-scale dataset for long text and multi-table summarization.
2 papers · 0 benchmarks
GameWikiSum is a domain-specific (video game) dataset for multi-document summarization, which is one hundred times larger than commonly used datasets, and in another domain than news.
1 paper · 0 benchmarks
HowSumm is a large-scale query-focused multi-document summarization dataset.
1 paper · 2 benchmarks
PMC-SA (PMC Structured Abstracts)
PMC-SA (PMC Structured Abstracts) is a dataset of academic publications, used for the task of structured summarization.
1 paper · 0 benchmarks
A high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization.
1 paper · 0 benchmarks
SubSumE Dataset This repository contains the SubSumE dataset for subjective document summarization.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.