{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/escape-a-large-scale-synthetic-corpus-for","title":"eSCAPE: a Large-scale Synthetic Corpus for Automatic Post-Editing","arxiv_id":"1803.07274","date":"2018-03-20","proceeding":"LREC 2018 5","authors":["Matteo Negri","Marco Turchi","Rajen Chatterjee","Nicola Bertoldi"],"abstract":"Training models for the automatic correction of machine-translated text\nusually relies on data consisting of (source, MT, human post- edit) triplets\nproviding, for each source sentence, examples of translation errors with the\ncorresponding corrections made by a human post-editor. Ideally, a large amount\nof data of this kind should allow the model to learn reliable correction\npatterns and effectively apply them at test stage on unseen (source, MT) pairs.\nIn practice, however, their limited availability calls for solutions that also\nintegrate in the training process other sources of knowledge. Along this\ndirection, state-of-the-art results have been recently achieved by systems\nthat, in addition to a limited amount of available training data, exploit\nartificial corpora that approximate elements of the \"gold\" training instances\nwith automatic translations. Following this idea, we present eSCAPE, the\nlargest freely-available Synthetic Corpus for Automatic Post-Editing released\nso far. eSCAPE consists of millions of entries in which the MT element of the\ntraining triplets has been obtained by translating the source side of\npublicly-available parallel corpora, and using the target side as an artificial\nhuman post-edit. Translations are obtained both with phrase-based and neural\nmodels. For each MT paradigm, eSCAPE contains 7.2 million triplets for\nEnglish-German and 3.3 millions for English-Italian, resulting in a total of\n14,4 and 6,6 million instances respectively. The usefulness of eSCAPE is proved\nthrough experiments in a general-domain scenario, the most challenging one for\nautomatic post-editing. For both language directions, the models trained on our\nartificial data always improve MT quality with statistically significant gains.\nThe current version of eSCAPE can be freely downloaded from:\nhttp://hltshare.fbk.eu/QT21/eSCAPE.html.","url_abs":"http://arxiv.org/abs/1803.07274v1","url_pdf":"http://arxiv.org/pdf/1803.07274v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"automatic-post-editing","task_name":"Automatic Post-Editing"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[{"slug":"escape","name":"eSCAPE","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.07274","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}