{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/quantitative-fine-grained-human-evaluation-of","title":"Quantitative Fine-Grained Human Evaluation of Machine Translation Systems: a Case Study on English to Croatian","arxiv_id":"1802.01451","date":"2018-02-02","proceeding":null,"authors":["Filip Klubička","Antonio Toral","Víctor M. Sánchez-Cartagena"],"abstract":"This paper presents a quantitative fine-grained manual evaluation approach to\ncomparing the performance of different machine translation (MT) systems. We\nbuild upon the well-established Multidimensional Quality Metrics (MQM) error\ntaxonomy and implement a novel method that assesses whether the differences in\nperformance for MQM error types between different MT systems are statistically\nsignificant. We conduct a case study for English-to-Croatian, a language\ndirection that involves translating into a morphologically rich language, for\nwhich we compare three MT systems belonging to different paradigms: pure\nphrase-based, factored phrase-based and neural. First, we design an\nMQM-compliant error taxonomy tailored to the relevant linguistic phenomena of\nSlavic languages, which made the annotation process feasible and accurate.\nErrors in MT outputs were then annotated by two annotators following this\ntaxonomy. Subsequently, we carried out a statistical analysis which showed that\nthe best-performing system (neural) reduces the errors produced by the worst\nsystem (pure phrase-based) by more than half (54\\%). Moreover, we conducted an\nadditional analysis of agreement errors in which we distinguished between short\n(phrase-level) and long distance (sentence-level) errors. We discovered that\nphrase-based MT approaches are of limited use for long distance agreement\nphenomena, for which neural MT was found to be especially effective.","url_abs":"http://arxiv.org/abs/1802.01451v1","url_pdf":"http://arxiv.org/pdf/1802.01451v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"quantitative-fine-grained-human-evaluation-of","repo_url":"https://github.com/GreenParachute/mqm-eng-cro","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}