{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ruber-an-unsupervised-method-for-automatic","title":"RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems","arxiv_id":"1701.03079","date":"2017-01-11","proceeding":null,"authors":["Chongyang Tao","Lili Mou","Dongyan Zhao","Rui Yan"],"abstract":"Open-domain human-computer conversation has been attracting increasing\nattention over the past few years. However, there does not exist a standard\nautomatic evaluation metric for open-domain dialog systems; researchers usually\nresort to human annotation for model evaluation, which is time- and\nlabor-intensive. In this paper, we propose RUBER, a Referenced metric and\nUnreferenced metric Blended Evaluation Routine, which evaluates a reply by\ntaking into consideration both a groundtruth reply and a query (previous\nuser-issued utterance). Our metric is learnable, but its training does not\nrequire labels of human satisfaction. Hence, RUBER is flexible and extensible\nto different datasets and languages. Experiments on both retrieval and\ngenerative dialog systems show that RUBER has a high correlation with human\nannotation.","url_abs":"http://arxiv.org/abs/1701.03079v2","url_pdf":"http://arxiv.org/pdf/1701.03079v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ruber-an-unsupervised-method-for-automatic","repo_url":"https://github.com/thu-coai/OpenMEVA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"dialogue-evaluation","task_name":"Dialogue Evaluation"},{"task_slug":"open-domain-dialog","task_name":"Open-Domain Dialog"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.03079","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}