{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-simple-recipe-for-multilingual-grammatical","title":"A Simple Recipe for Multilingual Grammatical Error Correction","arxiv_id":"2106.03830","date":"2021-06-07","proceeding":"ACL 2021 5","authors":["Sascha Rothe","Jonathan Mallinson","Eric Malmi","Sebastian Krause","Aliaksei Severyn"],"abstract":"This paper presents a simple recipe to train state-of-the-art multilingual Grammatical Error Correction (GEC) models. We achieve this by first proposing a language-agnostic method to generate a large number of synthetic examples. The second ingredient is to use large-scale multilingual language models (up to 11B parameters). Once fine-tuned on language-specific supervised sets we surpass the previous state-of-the-art results on GEC benchmarks in four languages: English, Czech, German and Russian. Having established a new set of baselines for GEC, we make our results easily reproducible and accessible by releasing a cLang-8 dataset. It is produced by using our best model, which we call gT5, to clean the targets of a widely used yet noisy lang-8 dataset. cLang-8 greatly simplifies typical GEC training pipelines composed of multiple fine-tuning stages -- we demonstrate that performing a single fine-tuning step on cLang-8 with the off-the-shelf language models yields further accuracy improvements over an already top-performing gT5 model for English.","url_abs":"https://arxiv.org/abs/2106.03830v2","url_pdf":"https://arxiv.org/pdf/2106.03830v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-simple-recipe-for-multilingual-grammatical","repo_url":"https://github.com/google-research-datasets/clang8","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-simple-recipe-for-multilingual-grammatical","repo_url":"https://github.com/gotutiyan/gec-t5","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"grammatical-error-correction","task_name":"Grammatical Error Correction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/grammatical-error-correction-on-conll-2014","task":"Grammatical Error Correction","dataset":"CoNLL-2014 Shared Task","model":"T5","rank_in_archive_order":7,"of":23,"metrics":{"F0.5":"68.87"},"uses_additional_data":false},{"leaderboard":"/sota/grammatical-error-correction-on-falko-merlin","task":"Grammatical Error Correction","dataset":"Falko-MERLIN","model":"gT5 xxl","rank_in_archive_order":3,"of":6,"metrics":{"F0.5":"75.96"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2106.03830","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}