Datasets › FCE
FCE (First Certificate in English)
The Cambridge Learner Corpus First Certificate in English (CLC FCE) dataset consists of short texts, written by learners of English as an additional language in response to exam prompts eliciting free-text answers and assessing mastery of the upper-intermediate proficiency level. The texts have been manually error-annotated using a taxonomy of 77 error types. The full dataset consists of 323,192 sentences. The publicly released subset of the dataset, named FCE-public, consists of 33,673 sentences split into test and training sets of 2,720 and 30,953 sentences, respectively.
Source: Compositional Sequence Labeling Models for Error Detection in Learner Writing
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Grammatical Error Detection | FCE | VERNet F0.5 72.2 | Neural Quality Estimation with Multiple Hypotheses for... | thunlp/VERNet | 8 | Compare |
Papers archive 2025-07-28
8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 151. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Neural Quality Estimation with Multiple Hypotheses for Grammatical Error Correction | 1 | 1 | 10 May 2021 | not harvested |
| Jointly Learning to Label Sentences and Tokens | 2 | 1 | 14 Nov 2018 | not harvested |
| Grammatical Error Detection Using Error- and Grammaticality-Specific Word Embeddings | 1 | 1 | 1 Nov 2017 | not harvested |
| Auxiliary Objectives for Neural Error Detection Models | 0 | 1 | 17 Jul 2017 | not harvested |
| Artificial Error Generation with Machine Translation and Syntactic Patterns | 0 | 1 | 17 Jul 2017 | not harvested |
| Semi-supervised Multitask Learning for Sequence Labeling | 3 | 1 | 24 Apr 2017 | not harvested |
| Attending to Characters in Neural Sequence Labeling Models | 0 | 1 | 14 Nov 2016 | not harvested |
| Compositional Sequence Labeling Models for Error Detection in Learner Writing | 0 | 1 | 20 Jul 2016 | not harvested |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- FCE
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections