{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/revisiting-summarization-evaluation-for","title":"Revisiting Summarization Evaluation for Scientific Articles","arxiv_id":"1604.00400","date":"2016-04-01","proceeding":"LREC 2016 5","authors":["Arman Cohan","Nazli Goharian"],"abstract":"Evaluation of text summarization approaches have been mostly based on metrics\nthat measure similarities of system generated summaries with a set of human\nwritten gold-standard summaries. The most widely used metric in summarization\nevaluation has been the ROUGE family. ROUGE solely relies on lexical overlaps\nbetween the terms and phrases in the sentences; therefore, in cases of\nterminology variations and paraphrasing, ROUGE is not as effective. Scientific\narticle summarization is one such case that is different from general domain\nsummarization (e.g. newswire data). We provide an extensive analysis of ROUGE's\neffectiveness as an evaluation metric for scientific summarization; we show\nthat, contrary to the common belief, ROUGE is not much reliable in evaluating\nscientific summaries. We furthermore show how different variants of ROUGE\nresult in very different correlations with the manual Pyramid scores. Finally,\nwe propose an alternative metric for summarization evaluation which is based on\nthe content relevance between a system generated summary and the corresponding\nhuman written summaries. We call our metric SERA (Summarization Evaluation by\nRelevance Analysis). Unlike ROUGE, SERA consistently achieves high correlations\nwith manual scores which shows its effectiveness in evaluation of scientific\narticle summarization.","url_abs":"http://arxiv.org/abs/1604.00400v1","url_pdf":"http://arxiv.org/pdf/1604.00400v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"revisiting-summarization-evaluation-for","repo_url":"https://github.com/jessicalopezespejel/gesera","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"text-summarization","task_name":"Text Summarization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.00400","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}