Papers › Stronger Baselines for Grammatical Error Correction Using Pretrained Encoder-Decoder Model

Stronger Baselines for Grammatical Error Correction Using Pretrained Encoder-Decoder Model

24 May 2020arXiv:2005.11849archive 2025-07-28

Satoru Katsumata, Mamoru Komachi

Studies on grammatical error correction (GEC) have reported the effectiveness of pretraining a Seq2Seq model with a large amount of pseudodata. However, this approach requires time-consuming pretraining for GEC because of the size of the pseudodata. In this study, we explore the utility of bidirectional and auto-regressive transformers (BART) as a generic pretrained encoder-decoder model for GEC. With the use of this generic pretrained model for GEC, the time-consuming pretraining can be eliminated. We find that monolingual and multilingual BART models achieve high performance in GEC, with one of the results being comparable to the current strong results in English GEC. Our implementations are publicly available at GitHub (https://github.com/Katsumata420/generic-pretrained-GEC).

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Katsumata420/generic-pretrained-GEC officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderGrammatical Error Correction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Grammatical Error Correction CoNLL-2014 Shared Task BART F0.5 63.0 #16 of 23 Archive leaderboard report
Grammatical Error Correction CoNLL-2014 Shared Task BART Precision 69.9 #16 of 23 Archive leaderboard report
Grammatical Error Correction CoNLL-2014 Shared Task BART Recall 45.1 #16 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionBARTBPEDense ConnectionsDropoutLSTMLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSeq2SeqSigmoid ActivationSoftmaxTanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections