{"url":"/dataset/summeval","name":"SummEval","full_name":null,"description_markdown":"The SummEval dataset is a resource developed by the Yale LILY Lab and Salesforce Research for evaluating text summarization models. It was created as part of a project to address shortcomings in summarization evaluation methods.\r\n\r\nThe dataset includes summaries generated by various recent summarization models trained on the CNN/DailyMail dataset. It also contains human annotations, collected from both crowdsource workers and experts. However, the source articles used to generate the summaries are not included.\r\n\r\nThe SummEval project also provides a toolkit for summarization evaluation. This toolkit unifies metrics and promotes robust comparison of summarization systems. It contains popular and recent metrics for summarization as well as several machine translation metrics.\r\n\r\nThe goal of the SummEval project is to promote a more complete evaluation protocol for text summarization and advance research in developing evaluation metrics that better correlate with human judgments.","description_withheld":null,"homepage":"https://github.com/Yale-LILY/SummEval","introduced_date":"2020-07-24","introduced_date_note":null,"introduced_by":{"paper":"/paper/summeval-re-evaluating-summarization","title":"SummEval: Re-evaluating Summarization Evaluation","first_author":"Alexander R. Fabbri","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["SummEval"],"data_loaders":[],"num_papers_in_archive":141,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}