{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-standardizing-korean-grammatical","title":"Towards standardizing Korean Grammatical Error Correction: Datasets and Annotation","arxiv_id":"2210.14389","date":"2022-10-25","proceeding":null,"authors":["Soyoung Yoon","Sungjoon Park","Gyuwan Kim","Junhee Cho","Kihyo Park","Gyutae Kim","Minjoon Seo","Alice Oh"],"abstract":"Research on Korean grammatical error correction (GEC) is limited, compared to other major languages such as English. We attribute this problematic circumstance to the lack of a carefully designed evaluation benchmark for Korean GEC. In this work, we collect three datasets from different sources (Kor-Lang8, Kor-Native, and Kor-Learner) that covers a wide range of Korean grammatical errors. Considering the nature of Korean grammar, We then define 14 error types for Korean and provide KAGAS (Korean Automatic Grammatical error Annotation System), which can automatically annotate error types from parallel corpora. We use KAGAS on our datasets to make an evaluation benchmark for Korean, and present baseline models trained from our datasets. We show that the model trained with our datasets significantly outperforms the currently used statistical Korean GEC system (Hanspell) on a wider range of error types, demonstrating the diversity and usefulness of the datasets. The implementations and datasets are open-sourced.","url_abs":"https://arxiv.org/abs/2210.14389v3","url_pdf":"https://arxiv.org/pdf/2210.14389v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-standardizing-korean-grammatical","repo_url":"https://github.com/soyoung97/standard_korean_gec","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"grammatical-error-correction","task_name":"Grammatical Error Correction"}],"methods":[],"datasets_introduced":[{"slug":"kor-lang8","name":"Kor-Lang8","full_name":"Lang-8 Korean Corpus"},{"slug":"kor-learner","name":"Kor-Learner","full_name":"Korean Learner Corpus"},{"slug":"kor-native","name":"Kor-Native","full_name":"Native Korean Corpus"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.14389","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.14389"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/soyoung97/standard_korean_gec","reach":null}],"summary":{"ran_draft_wrong":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"6160cddabde132fe","entry":"formatEdit","repo":"soyoung97/standard_korean_gec","repo_kind":"official","path":"KAGAS/parallel_to_m2_korean.py","file_url":"https://github.com/soyoung97/standard_korean_gec/blob/HEAD/KAGAS/parallel_to_m2_korean.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6160cddabde132fe"}},{"code_sha256_prefix":"c1ace94cc6c43b23","entry":"get_gleu_stats","repo":"soyoung97/standard_korean_gec","repo_kind":"official","path":"eval/run_gleu.py","file_url":"https://github.com/soyoung97/standard_korean_gec/blob/HEAD/eval/run_gleu.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c1ace94cc6c43b23"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}