{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-large-scale-test-set-for-the-evaluation-of","title":"A Large-Scale Test Set for the Evaluation of Context-Aware Pronoun Translation in Neural Machine Translation","arxiv_id":"1810.02268","date":"2018-10-04","proceeding":"WS 2018 10","authors":["Mathias Müller","Annette Rios","Elena Voita","Rico Sennrich"],"abstract":"The translation of pronouns presents a special challenge to machine\ntranslation to this day, since it often requires context outside the current\nsentence. Recent work on models that have access to information across sentence\nboundaries has seen only moderate improvements in terms of automatic evaluation\nmetrics such as BLEU. However, metrics that quantify the overall translation\nquality are ill-equipped to measure gains from additional context. We argue\nthat a different kind of evaluation is needed to assess how well models\ntranslate inter-sentential phenomena such as pronouns. This paper therefore\npresents a test suite of contrastive translations focused specifically on the\ntranslation of pronouns. Furthermore, we perform experiments with several\ncontext-aware models. We show that, while gains in BLEU are moderate for those\nsystems, they outperform baselines by a large margin in terms of accuracy on\nour contrastive test set. Our experiments also show the effectiveness of\nparameter tying for multi-encoder architectures.","url_abs":"http://arxiv.org/abs/1810.02268v3","url_pdf":"http://arxiv.org/pdf/1810.02268v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-large-scale-test-set-for-the-evaluation-of","repo_url":"https://github.com/ZurichNLP/ContraPro","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1810.02268","atlas_url":"https://app.syntology.ai/?focus=1810.02268","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}