Datasets › SummEval

SummEval

Introduced by Alexander R. Fabbri et al. in SummEval: Re-evaluating Summarization Evaluation24 Jul 2020 archive 2025-07-28

The SummEval dataset is a resource developed by the Yale LILY Lab and Salesforce Research for evaluating text summarization models. It was created as part of a project to address shortcomings in summarization evaluation methods.

The dataset includes summaries generated by various recent summarization models trained on the CNN/DailyMail dataset. It also contains human annotations, collected from both crowdsource workers and experts. However, the source articles used to generate the summaries are not included.

The SummEval project also provides a toolkit for summarization evaluation. This toolkit unifies metrics and promotes robust comparison of summarization systems. It contains popular and recent metrics for summarization as well as several machine translation metrics.

The goal of the SummEval project is to promote a more complete evaluation protocol for text summarization and advance research in developing evaluation metrics that better correlate with human judgments.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 141 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • SummEval

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections