Browse State-of-the-Art › Spelling Correction
Spelling Correction
52 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Spelling correction is the task of detecting and correcting spelling mistakes.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 52 papers with code (193 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Jul 2016 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe present an approach to training neural networks to generate sequences using actor-critic methods from reinforcement learning (RL).
-
18 Aug 2023 2 repositories listedOur research mainly focuses on exploring natural spelling errors and mistypings in texts and studying the ways those errors can be emulated in correct sentences to effectively enrich generative models' pre-train…
-
17 Aug 2023 2 repositories listedHowever, we note a critical flaw in the process of tagging one character to another, that the correction is excessively conditioned on the error.
-
12 Feb 2023 2 repositories listedWe extend a current sequence-tagging approach to Grammatical Error Correction (GEC) by introducing specialised tags for spelling correction and morphological inflection using the SymSpell and LemmInflect algorithms.
-
15 Oct 2020 2 repositories listedWe identify three key ingredients of high-quality tokenization repair, all missing from previous work: deep language models with a bidirectional component, training the models on text with spelling errors, and making…
-
1 Jul 2019 2 repositories listedThere are a lot of noise texts surrounding a person in modern life.
-
10 Oct 2017 2 repositories listedWe show that MoNoise beats the state-of-the-art on different normalization benchmarks for English and Dutch, which all define the task of normalization slightly different.
-
12 May 2025 1 repository listedTo tackle this, we propose a data augmentation approach using unlabeled text to generate multi-level corruptions, and introduce TiSpell, a semi-masked model capable of correcting both character- and syllable-level…
-
10 Apr 2025 1 repository listedHowever, LLMs face limitations in CSC, particularly over-correction, making them suboptimal for this task.
-
21 Feb 2025 1 repository listedTo address this issue, we introduce the task of General Chinese Character Error Correction (C2EC), which focuses on all three types of character errors.
-
26 Nov 2024 1 repository listedTokenization methods like Byte-Pair Encoding (BPE) enhance computational efficiency in large language models (LLMs) but often obscure internal character structures within tokens.
-
18 Nov 2024 1 repository listedThe task of converting Hanyu Pinyin abbreviations to Chinese characters is a significant branch within the domain of Chinese Spelling Correction (CSC).
-
5 Oct 2024 1 repository listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)This work proposes a simple training-free prompt-free approach to leverage large language models (LLMs) for the Chinese spelling correction (CSC) task, which is totally different from all previous CSC approaches.
-
8 Sep 2024 1 repository listedChinese Spelling Correction (CSC) aims to detect and correct spelling errors in Chinese sentences caused by phonetic or visual similarities.
-
6 Sep 2024 1 repository listedOur detector is designed to yield two error detection results, each characterized by high precision and recall.
-
11 May 2024 1 repository listedSpelling correction is the task of identifying spelling mistakes, typos, and grammatical mistakes in a given text and correcting them according to their context and grammatical structure.
-
14 Nov 2023 1 repository listedHowever, in the Chinese Spelling Correction (CSC) task, we observe a discrepancy: while ChatGPT performs well under human evaluation, it scores poorly according to traditional metrics.
-
20 Jun 2023 1 repository listedOur contribution is in showing that it can be made highly scalable through a simple relaxation of the objective and a highly efficient implementation.
-
4 Jun 2023 1 repository listedContextual spelling correction models are an alternative to shallow fusion to improve automatic speech recognition (ASR) quality given user vocabulary.
-
28 May 2023 1 repository listedIn this paper, we study Chinese Spelling Correction (CSC) as a joint decision made by two separate models: a language model and an error model.
-
24 May 2023 1 repository listedChinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts.
-
16 Jan 2023 1 repository listedBy borrowing the powerful ability of BERT, we propose a novel zero-shot error detection method to do a preliminary detection, which guides our model to attend more on the probably wrong tokens in encoding and to avoid…
-
19 Dec 2022 1 repository listedLanguage tasks involving character-level manipulations (e.
-
16 Nov 2022 1 repository listedIn this paper, we present CSCD-NS, the first Chinese spelling check (CSC) dataset designed for native speakers, containing 40, 000 samples from a Chinese social platform.
-
21 Oct 2022 1 repository listedIn this work, we define the task of Medical-domain Chinese Spelling Correction and propose MCSCSet, a large scale specialist-annotated dataset that contains about 200k samples.
-
6 Oct 2022 1 repository listedWith 84.
-
20 Aug 2022 1 repository listedA specialized BERT model named BSpell has been proposed in this paper targeted towards word for word correction in sentence level.
-
8 Jul 2022 1 repository listedAbbreviations and contractions are commonly found in text across different domains.
-
1 May 2022 1 repository listedThese methods have two limitations: (1) they have poor performance on multi-typo texts.
-
2 Mar 2022 1 repository listedIn this work, we introduce a novel approach to do contextual biasing by adding a contextual spelling correction model on top of the end-to-end ASR system.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections