Papers › System Combination via Quality Estimation for Grammatical Error Correction

System Combination via Quality Estimation for Grammatical Error Correction

23 Oct 2023arXiv:2310.14947archive 2025-07-28

Muhammad Reza Qorib, Hwee Tou Ng

Quality estimation models have been developed to assess the corrections made by grammatical error correction (GEC) models when the reference or gold-standard corrections are not available. An ideal quality estimator can be utilized to combine the outputs of multiple GEC systems by choosing the best subset of edits from the union of all edits proposed by the GEC base systems. However, we found that existing GEC quality estimation models are not good enough in differentiating good corrections from bad ones, resulting in a low F0.5 score when used for system combination. In this paper, we propose GRECO, a new state-of-the-art quality estimation model that gives a better estimate of the quality of a corrected sentence, as indicated by having a higher correlation to the F0.5 score of a corrected sentence. It results in a combined GEC system with a higher F0.5 score. We also propose three methods for utilizing GEC quality estimation models for system combination with varying generality: model-agnostic, model-agnostic with voting bias, and model-dependent method. The combined GEC system outperforms the state of the art on the CoNLL-2014 test set and the BEA-2019 test set, achieving the highest F0.5 scores published to date.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

nusnlp/greco officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Grammatical Error CorrectionSentence

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Grammatical Error Correction BEA-2019 (test) GRECO (voting+ESC) F0.5 80.84 #2 of 19 Archive leaderboard report
Grammatical Error Correction CoNLL-2014 Shared Task GRECO (voting+ESC) F0.5 71.12 #3 of 23 Archive leaderboard report
Grammatical Error Correction CoNLL-2014 Shared Task GRECO (voting+ESC) Precision 79.6 #3 of 23 Archive leaderboard report
Grammatical Error Correction CoNLL-2014 Shared Task GRECO (voting+ESC) Recall 49.86 #3 of 23 Archive leaderboard report
Grammatical Error Correction CoNLL-2014 Shared Task (10 annotations) GRECO (vote+ESC) F0.5 85.21 #1 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BASESET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections