{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rankme-reliable-human-ratings-for-natural","title":"RankME: Reliable Human Ratings for Natural Language Generation","arxiv_id":"1803.05928","date":"2018-03-15","proceeding":"NAACL 2018 6","authors":["Jekaterina Novikova","Ondřej Dušek","Verena Rieser"],"abstract":"Human evaluation for natural language generation (NLG) often suffers from\ninconsistent user ratings. While previous research tends to attribute this\nproblem to individual user preferences, we show that the quality of human\njudgements can also be improved by experimental design. We present a novel\nrank-based magnitude estimation method (RankME), which combines the use of\ncontinuous scales and relative assessments. We show that RankME significantly\nimproves the reliability and consistency of human ratings compared to\ntraditional evaluation methods. In addition, we show that it is possible to\nevaluate NLG systems according to multiple, distinct criteria, which is\nimportant for error analysis. Finally, we demonstrate that RankME, in\ncombination with Bayesian estimation of system quality, is a cost-effective\nalternative for ranking multiple NLG systems.","url_abs":"http://arxiv.org/abs/1803.05928v1","url_pdf":"http://arxiv.org/pdf/1803.05928v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rankme-reliable-human-ratings-for-natural","repo_url":"https://github.com/jeknov/RankME","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"experimental-design","task_name":"Experimental Design"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.05928","atlas_url":"https://app.syntology.ai/?focus=1803.05928","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}