Home › Datasets › task › Grammatical Error Detection

Grammatical Error Detection datasets

archive 2025-07-28

6 datasets carry the task tag "Grammatical Error Detection" (the task itself: Grammatical Error Detection), ordered by the archive's paper count. Page 1 of 1: 6 shown of 6. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Grammatical Error Detection datasets 1–6 of 6

The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
FCE (First Certificate in English)
The Cambridge Learner Corpus First Certificate in English (CLC FCE) dataset consists of short texts, written by learners of English as an additional language in response to exam prompts eliciting free-text answers and assessing mastery of…
151 papers · 1 benchmark
JFLEG (JHU FLuency-Extended GUG corpus)
JFLEG is for developing and evaluating grammatical error correction (GEC).
91 papers · 5 benchmarks
Lang-8 Preprocessed Dataset (for GED): - Dataset: Lang-8, a publicly available dataset containing user-generated content, primarily from second-language learners, focused on writing errors.
1 paper · 0 benchmarks
FCGEC (FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction)
a fine-grained corpus to detect, identify and correct the chinese grammatical errors.
1 paper · 1 benchmark
We present a new annotated corpus of written learner English, derived from essays submitted to the learning platform Write & Improve (W&I).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.