Papers › To Err Is Human, but Llamas Can Learn It Too

To Err Is Human, but Llamas Can Learn It Too

8 Mar 2024arXiv:2403.05493archive 2025-07-28

Agnes Luhtaru, Taido Purason, Martin Vainikko, Maksym Del, Mark Fishel

This study explores enhancing grammatical error correction (GEC) through artificial error generation (AEG) using language models (LMs). Specifically, we fine-tune Llama 2-based LMs for error generation and find that this approach yields synthetic errors akin to human errors. Next, we train GEC Llama models with the help of these artificial errors and outperform previous state-of-the-art error correction models, with gains ranging between 0.8 and 6 F0.5 points across all tested languages (German, Ukrainian, and Estonian). Moreover, we demonstrate that generating errors by fine-tuning smaller sequence-to-sequence models and prompting large commercial LMs (GPT-3.5 and GPT-4) also results in synthetic errors beneficially affecting error generation models.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

TartuNLP/gec-llm officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Grammatical Error Correction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Grammatical Error Correction EstGEC-L2 Llama + 1M BT + gold F0.5 69.97 #1 of 1 Archive leaderboard report
Grammatical Error Correction Falko-MERLIN Llama + 1M BT + gold F0.5 76.75 #1 of 6 Archive leaderboard report
Grammatical Error Correction UA-GEC Llama + 1M BT + gold F0.5 74.09 #1 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections