Home › Datasets › task › Text Summarization
Text Summarization datasets
archive 2025-07-28
98 datasets carry the task tag "Text Summarization" (the task itself: Text Summarization), ordered by the archive's paper count. Page 1 of 3: 48 shown of 98. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Text Summarization datasets 1–48 of 98
The PubMed dataset consists of 19717 scientific publications from PubMed database pertaining to diabetes classified into one of three classes.
1,236 papers · 19 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
CNN/Daily Mail is a dataset for text summarization.
530 papers · 8 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
MTEB (Massive Text Embedding Benchmark)
MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages.
155 papers · 6 benchmarks
A large corpus of 81.1M English-language academic papers spanning many academic disciplines.
152 papers · 1 benchmark
A new dataset with abstractive dialogue summaries.
145 papers · 4 benchmarks
WikiHow is a dataset of more than 230,000 article and summary pairs extracted and constructed from an online knowledge base written by different human authors.
127 papers · 2 benchmarks
Multi-News, consists of news articles and human-written summaries of these articles from the site newser.com.
122 papers · 5 benchmarks
CORNELL NEWSROOM is a large dataset for training and evaluating summarization systems.
107 papers · 0 benchmarks
KP20k is a large-scale scholarly articles dataset with 528K articles for training, 20K articles for validation and 20K articles for testing.
87 papers · 3 benchmarks
The How2 dataset contains 13,500 videos, or 300 hours of speech, and is split into 185,187 training, 2022 development (dev), and 2361 test utterances.
84 papers · 2 benchmarks
GovReport is a dataset for long document summarization, with significantly longer documents and summaries.
83 papers · 2 benchmarks
QMSum is a new human-annotated benchmark for query-based multi-domain meeting summarisation task, which consists of 1,808 query-summary pairs over 232 meetings in multiple domains.
69 papers · 1 benchmark
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
Sentence Compression is a dataset where the syntactic trees of the compressions are subtrees of their uncompressed counterparts, and hence where supervised systems which require a structural alignment between the input and output can be…
63 papers · 0 benchmarks
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.
62 papers · 2 benchmarks
LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public.
58 papers · 2 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
Consists of 1.3 million records of U.S.
50 papers · 2 benchmarks
BillSum is the first dataset for summarization of US Congressional and California state bills.
47 papers · 2 benchmarks
Reddit TIFU dataset is a newly collected Reddit dataset, where TIFU denotes the name of /r/tifu subbreddit.
47 papers · 1 benchmark
MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
SCROLLS (Standardized CompaRison Over Long Language Sequences)
SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts.
42 papers · 1 benchmark
BookSum is a collection of datasets for long-form narrative summarization.
39 papers · 2 benchmarks
Arxiv HEP-TH (high energy physics theory) citation graph is from the e-print arXiv and covers all the citations within a dataset of 27,770 papers with 352,807 edges.
35 papers · 5 benchmarks
xP3 is a multilingual dataset for multitask prompted finetuning.
34 papers · 0 benchmarks
MeQSum is a dataset for medical question summarization.
33 papers · 1 benchmark
The Extreme Summarization (XSum) dataset is a dataset for evaluation of abstractive single-document summarization systems.
32 papers · 5 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
To study the task of email subject line generation: automatically generating an email subject line from the email body.
22 papers · 1 benchmark
Large-scale manually-annotated corpus for 1,000 scientific papers (on computational linguistics) for automatic summarization.
20 papers · 0 benchmarks
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S.
16 papers · 1 benchmark
The DUC2004 dataset is a dataset for document summarization.
15 papers · 4 benchmarks
KorSTS is a dataset for semantic textural similarity (STS) in Korean.
13 papers · 0 benchmarks
This is a dataset for evaluating summarisation methods for research papers.
13 papers · 3 benchmarks
ECTSum is a dataset with transcripts of earnings calls (ECTs), hosted by public companies, as documents, and short experts-written telegram-style bullet point summaries derived from corresponding Reuters articles.
12 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation
8 papers · 1 benchmark
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks
Source: BARThez: a Skilled Pretrained French Sequence-to-Sequence Model OrangeSum is a single-document extreme summarization dataset with two tasks: title and abstract.
8 papers · 1 benchmark
WikiCatSum is a domain specific Multi-Document Summarisation (MDS) dataset.
7 papers · 0 benchmarks
This dataset was created using a dataset used for data categorization that onsists of 2225 documents from the BBC news website corresponding to stories in five topical areas from 2004-2005 used in the paper of D.
6 papers · 0 benchmarks
CELLS is a large (63k pairs) and broadest-ranging (12 journals) parallel corpus for lay language generation.
6 papers · 0 benchmarks
ConvoSumm is a suite of four datasets to evaluate a model’s performance on a broad spectrum of conversation data.
6 papers · 0 benchmarks
The TalkSumm dataset contains 1705 automatically-generated summaries of scientific papers from ACL, NAACL, EMNLP, SIGDIAL (2015-2018), and ICML (2017-2018).
6 papers · 0 benchmarks
Klexikon (Klexikon: A German Dataset for Joint Summarization and Simplification)
The dataset introduces document alignments between German Wikipedia and the children's lexicon Klexikon.
5 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.