{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-price-of-debiasing-automatic-metrics-in","title":"The price of debiasing automatic metrics in natural language evaluation","arxiv_id":"1807.02202","date":"2018-07-06","proceeding":null,"authors":["Arun Tejasvi Chaganty","Stephen Mussman","Percy Liang"],"abstract":"For evaluating generation systems, automatic metrics such as BLEU cost\nnothing to run but have been shown to correlate poorly with human judgment,\nleading to systematic bias against certain model improvements. On the other\nhand, averaging human judgments, the unbiased gold standard, is often too\nexpensive. In this paper, we use control variates to combine automatic metrics\nwith human evaluation to obtain an unbiased estimator with lower cost than\nhuman evaluation alone. In practice, however, we obtain only a 7-13% cost\nreduction on evaluating summarization and open-response question answering\nsystems. We then prove that our estimator is optimal: there is no unbiased\nestimator with lower cost. Our theory further highlights the two fundamental\nbottlenecks---the automatic metric and the prompt shown to human\nevaluators---both of which need to be improved to obtain greater cost savings.","url_abs":"http://arxiv.org/abs/1807.02202v1","url_pdf":"http://arxiv.org/pdf/1807.02202v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-price-of-debiasing-automatic-metrics-in","repo_url":"https://worksheets.codalab.org/worksheets/0xbda93e6519134c1ab1893ceaa19c8a5c","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.02202","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}