Home › Datasets › task › Grammatical Error Correction
Grammatical Error Correction datasets
archive 2025-07-28
16 datasets carry the task tag "Grammatical Error Correction" (the task itself: Grammatical Error Correction), ordered by the archive's paper count. Page 1 of 1: 16 shown of 16. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Grammatical Error Correction datasets 1–16 of 16
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
JFLEG (JHU FLuency-Extended GUG corpus)
JFLEG is for developing and evaluating grammatical error correction (GEC).
91 papers · 5 benchmarks
CoNLL-2014 will continue the CoNLL tradition of having a high profile shared task in natural language processing.
79 papers · 0 benchmarks
WI-LOCNESS (Cambridge English Write & Improve & LOCNESS)
WI-LOCNESS is part of the Building Educational Applications 2019 Shared Task for Grammatical Error Correction.
29 papers · 2 benchmarks
MuCGEC (Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction)
MuCGEC is a multi-reference multi-source evaluation dataset for Chinese Grammatical Error Correction (CGEC), consisting of 7,063 sentences collected from three different Chinese-as-a-Second-Language (CSL) learner sources.
15 papers · 1 benchmark
AKCES-GEC is a new dataset on grammatical error correction for Czech.
11 papers · 0 benchmarks
UA-GEC (UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language)
UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language
9 papers · 1 benchmark
Grammatical error correction dataset for text from Wikipedia.
4 papers · 0 benchmarks
Are you the kind of person who makes a lot of typos when writing code?
3 papers · 0 benchmarks
Grammatical error correction dataset for text from Yahoo!
2 papers · 0 benchmarks
FCGEC (FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction)
a fine-grained corpus to detect, identify and correct the chinese grammatical errors.
1 paper · 1 benchmark
Kor-Lang8 is a Korean grammatical error correction (GEC) dataset extracted from the NAIST Lang-8 Learner Corpora by the language label.
1 paper · 0 benchmarks
Kor-Learner is a Korean grammatical error correction (GEC) dataset made from the NIKL learner corpus containing essays written by Korean learners and their grammatical error correction annotations by their tutors in an morpheme-level XML…
1 paper · 0 benchmarks
Kor-Learner is a Korean grammatical error correction (GEC) dataset collected grammatically from two sources, and the correct sentences were read using Google Text-to-Speech(TTS) system.
1 paper · 0 benchmarks
NaSGEC is a new dataset to facilitate research on Chinese grammatical error correction (CGEC) for native speaker texts from multiple domains.
1 paper · 0 benchmarks
We present a new annotated corpus of written learner English, derived from essays submitted to the learning platform Write & Improve (W&I).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.