{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/summeval-re-evaluating-summarization","title":"SummEval: Re-evaluating Summarization Evaluation","arxiv_id":"2007.12626","date":"2020-07-24","proceeding":null,"authors":["Alexander R. Fabbri","Wojciech Kryściński","Bryan McCann","Caiming Xiong","Richard Socher","Dragomir Radev"],"abstract":"The scarcity of comprehensive up-to-date studies on evaluation metrics for text summarization and the lack of consensus regarding evaluation protocols continue to inhibit progress. We address the existing shortcomings of summarization evaluation methods along five dimensions: 1) we re-evaluate 14 automatic evaluation metrics in a comprehensive and consistent fashion using neural summarization model outputs along with expert and crowd-sourced human annotations, 2) we consistently benchmark 23 recent summarization models using the aforementioned automatic evaluation metrics, 3) we assemble the largest collection of summaries generated by models trained on the CNN/DailyMail news dataset and share it in a unified format, 4) we implement and share a toolkit that provides an extensible and unified API for evaluating summarization models across a broad range of automatic metrics, 5) we assemble and share the largest and most diverse, in terms of model types, collection of human judgments of model-generated summaries on the CNN/Daily Mail dataset annotated by both expert judges and crowd-source workers. We hope that this work will help promote a more complete evaluation protocol for text summarization as well as advance research in developing evaluation metrics that better correlate with human judgments.","url_abs":"https://arxiv.org/abs/2007.12626v4","url_pdf":"https://arxiv.org/pdf/2007.12626v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"summeval-re-evaluating-summarization","repo_url":"https://github.com/Yale-LILY/SummEval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"summeval-re-evaluating-summarization","repo_url":"https://github.com/PrimerAI/blanc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"summeval-re-evaluating-summarization","repo_url":"https://github.com/abhisha1991/w266_final_project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"summeval-re-evaluating-summarization","repo_url":"https://github.com/jungokasai/billboard","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"summeval-re-evaluating-summarization","repo_url":"https://github.com/jungokasai/thumb","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"summeval-re-evaluating-summarization","repo_url":"https://github.com/kzawisto/unused_information_llm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"text-summarization","task_name":"Text Summarization"}],"methods":[],"datasets_introduced":[{"slug":"summeval","name":"SummEval","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2007.12626","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2007.12626"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Yale-LILY/SummEval","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jungokasai/billboard","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kzawisto/unused_information_llm","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/PrimerAI/blanc","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/abhisha1991/w266_final_project","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jungokasai/thumb","reach":{"status":"ok"}}],"summary":{"unverified":6},"by_repo_kind":{"listed":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d1bb17547eab5079","entry":"batch_data","repo":"PrimerAI/blanc","repo_kind":"listed","path":"blanc/utils.py","file_url":"https://github.com/PrimerAI/blanc/blob/HEAD/blanc/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d1bb17547eab5079"}},{"code_sha256_prefix":"5e1adf0ee89abd21","entry":"get_model","repo":"PrimerAI/blanc","repo_kind":"listed","path":"blanc/shannon.py","file_url":"https://github.com/PrimerAI/blanc/blob/HEAD/blanc/shannon.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5e1adf0ee89abd21"}},{"code_sha256_prefix":"5915ef1c36c6c3f6","entry":"is_token_large_enough","repo":"PrimerAI/blanc","repo_kind":"listed","path":"blanc/utils.py","file_url":"https://github.com/PrimerAI/blanc/blob/HEAD/blanc/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5915ef1c36c6c3f6"}},{"code_sha256_prefix":"1ed41e9d17925038","entry":"load","repo":"PrimerAI/blanc","repo_kind":"listed","path":"shannon/summeval_score.py","file_url":"https://github.com/PrimerAI/blanc/blob/HEAD/shannon/summeval_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1ed41e9d17925038"}},{"code_sha256_prefix":"d61c81b9199510ad","entry":"mask_tokens_evenly","repo":"PrimerAI/blanc","repo_kind":"listed","path":"blanc/utils.py","file_url":"https://github.com/PrimerAI/blanc/blob/HEAD/blanc/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d61c81b9199510ad"}},{"code_sha256_prefix":"56f81e0b7efc8ed8","entry":"prepare_inputs_for_generation","repo":"PrimerAI/blanc","repo_kind":"listed","path":"blanc/shannon.py","file_url":"https://github.com/PrimerAI/blanc/blob/HEAD/blanc/shannon.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"56f81e0b7efc8ed8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}