Home › Datasets › task › Text Summarization

Text Summarization datasets

archive 2025-07-28

98 datasets carry the task tag "Text Summarization" (the task itself: Text Summarization), ordered by the archive's paper count. Page 2 of 3: 48 shown of 98. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Text Summarization datasets 49–96 of 98

A large-scale Indonesian summarization dataset consisting of harvested articles from Liputan6.com, an online news portal, resulting in 215,827 document-summary pairs.
5 papers · 0 benchmarks
SSN (Semantic Scholar Network)
SSN (short for Semantic Scholar Network) is a scientific papers summarization dataset which contains 141K research papers in different domains and 661K citation relationships.
5 papers · 0 benchmarks
WikiHowQA is a Community-based Question Answering dataset, which can be used for both answer selection and abstractive summarization tasks.
5 papers · 0 benchmarks
EurekaAlert (Eureka Alert)
This dataset contains around 5000 scholarly articles and their corresponding easy summary from eureka alert blog, the dataset can be used for the combined task of summarization and simplification.
4 papers · 2 benchmarks
Gazeta is a dataset for automatic summarization of Russian news.
4 papers · 1 benchmark
IndoNLG is a benchmark to measure natural language generation (NLG) progress in three low-resource—yet widely spoken—languages of Indonesia: Indonesian, Javanese, and Sundanese.
4 papers · 0 benchmarks
OASum is a large-scale open-domain aspect-based summarization dataset which contains more than 3.7 million instances with around 1 million different aspects on 2 million Wikipedia pages.
4 papers · 0 benchmarks
For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including…
4 papers · 1 benchmark
Pn-summary is a dataset for Persian abstractive text summarization.
4 papers · 0 benchmarks
CHQ-Summ (Consumer Healthcare Question Summarization)
Contains 1507 domain-expert annotated consumer health questions and corresponding summaries.
3 papers · 0 benchmarks
CNewSum is a large-scale Chinese news summarization dataset which consists of 304,307 documents and human-written summaries for the news feed.
3 papers · 0 benchmarks
A corpus of 553k news articles from six Persian news websites and agencies with relatively high quality author extracted keyphrases, which is then filtered and cleaned to achieve higher quality keyphrases.
3 papers · 0 benchmarks
PubMedCite is a domain-specific dataset with about 192K biomedical scientific papers and a large citation graph preserving 917K citation relationships between them.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
Given an English article, generate a short summary in the target language.
2 papers · 0 benchmarks
DUC 2007 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
2 papers · 0 benchmarks
DaNewsroom (DaNewsroom: A Large-scale Danish Summarisation Dataset)
The first large-scale non-English language dataset specifically curated for automatic summarisation.
2 papers · 0 benchmarks
An open corpus of Scientific Research papers which has a representative sample from across scientific disciplines.
2 papers · 0 benchmarks
FIB (Factual Inconsistency Benchmark)
Factual Inconsistency Benchmark (FIB) is a benchmark that focuses on the task of summarization.
2 papers · 0 benchmarks
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the…
2 papers · 0 benchmarks
MentSum (Mental Health Summarization Dataset)
Mental health remains a significant challenge of public health worldwide.
2 papers · 1 benchmark
OpenAsp Dataset OpenAsp is an Open Aspect-based Multi-Document Summarization dataset derived from DUC and MultiNews summarization datasets.
2 papers · 0 benchmarks
SuMe (A Dataset Towards Summarizing Biomedical Mechanisms)
Can language models read biomedical texts and explain the biomedical mechanisms discussed?
2 papers · 0 benchmarks
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a dataset for multi-document summarization in Portuguese, what means that it has examples of multiple documents (input) related to human-written summaries (output).
1 paper · 0 benchmarks
We present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396,209 papers.
1 paper · 0 benchmarks
ComSum is a data set of 7 million commit messages for text summarization.
1 paper · 0 benchmarks
DUC 2006 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
1 paper · 0 benchmarks
The "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies.
1 paper · 0 benchmarks
Inshorts News (Inshorts English News dataset)
Inshorts News dataset Inshorts provides a news summary in 60 words or less.
1 paper · 1 benchmark
MOS Dataset (Microblog Opinion Summarisation)
This dataset was used in the paper 'Template-based Abstractive Microblog Opinion Summarisation' (to be published at TACL, 2022).
1 paper · 0 benchmarks
MultiSum is a dataset for multimodal summarization (MSMO).
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks
PMC-SA (PMC Structured Abstracts)
PMC-SA (PMC Structured Abstracts) is a dataset of academic publications, used for the task of structured summarization.
1 paper · 0 benchmarks
This is a large-scale court judgment dataset, where each judgment is a summary of the case description with a patternized style.
1 paper · 0 benchmarks
Robust Summarization Evaluation Benchmark is a large human evaluation dataset consisting of over 22k summary-level annotations over state-of-the-art systems on three datasets.
1 paper · 0 benchmarks
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper: Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable…
1 paper · 0 benchmarks
SubSumE Dataset This repository contains the SubSumE dataset for subjective document summarization.
1 paper · 0 benchmarks
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
Bangladesh's legal system struggles with major challenges like delays, complexity, high costs, and millions of unresolved cases, which deter many from pursuing legal action due to lack of knowledge or financial constraints.
1 paper · 0 benchmarks
In doctor-patient conversations, identifying medically relevant information is crucial, posing the need for conversation summarization.
1 paper · 0 benchmarks
Wiki-en is an annotated English dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
Wiki-zh is an annotated Chinese dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
WikiDes is a dataset for generating descriptions of Wikidata from Wikipedia paragraphs.
1 paper · 0 benchmarks
WikiWeb2M (Wikipedia Webpage 2M)
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.