Browse State-of-the-Art › Lexical Normalization
Lexical Normalization
16 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Lexical normalization is the task of translating/transforming a non standard text to a standard register.
Example:
new pix comming tomoroe
new pictures coming tomorrow
Datasets usually consists of tweets, since these naturally contain a fair amount of these phenomena.
For lexical normalization, only replacements on the word-level are annotated. Some corpora include annotation for 1-N and N-1 replacements. However, word insertion/deletion and reordering is not part of the task.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LexNorm (4 rows) | MoNoise | MoNoise: Modeling Noise Using a Modular Normalization System | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
16 shown of 16 papers with code (47 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Oct 2017 2 repositories listedWe show that MoNoise beats the state-of-the-art on different normalization benchmarks for English and Dutch, which all define the task of normalization slightly different.
-
13 Jan 2025 1 repository listedViSoLex is an open-source system designed to address the unique challenges of lexical normalization for Vietnamese social media text.
-
29 Jan 2024 1 repository listedIn this work, we introduce Vietnamese Lexical Normalization (ViLexNorm), the first-ever corpus developed for the Vietnamese lexical normalization task.
-
12 Nov 2023 1 repository listedOur dataset is accessible for research purposes.
-
1 Oct 2022 1 repository listedAutomatically detecting the intent of an utterance is important for various downstream natural language processing tasks.
-
1 Nov 2021 1 repository listedThis task is beneficial for downstream analysis, as it provides a way to harmonize (often spontaneous) linguistic variation.
-
28 Oct 2021 1 repository listedWe present the winning entry to the Multilingual Lexical Normalization (MultiLexNorm) shared task at W-NUT 2021 (van der Goot et al., 2021a), which evaluates lexical-normalization systems on 12 social media datasets in…
-
24 May 2021 1 repository listedWe examine language-specific versus multilingual BERT, and study the effect of lexical normalization on NER.
-
8 Apr 2021 1 repository listedMorphological analysis (MA) and lexical normalization (LN) are both important tasks for Japanese user-generated text (UGT).
-
1 Apr 2021 1 repository listedLexical normalization, the translation of non-canonical data to standard language, has shown to improve the performance of many natural language processing tasks on social media.
-
31 Mar 2020 1 repository listedRoman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content.
-
4 Jan 2020 1 repository listedSuch informal and code-switched content are under-resourced in terms of labeled datasets and language models even for popular tasks like sentiment classification.
-
29 Nov 2019 1 repository listedOur model achieves high accuracy for classification on this dataset and outperforms the previous model for multilingual text classification, highlighting language independence of McM.
-
1 Jul 2019 1 repository listedIn this paper, we introduce and demonstrate the online demo as well as the command line interface of a lexical normalization system (MoNoise) for a variety of languages.
-
12 Apr 2019 1 repository listedSocial media offer an abundant source of valuable raw data, however informal writing can quickly become a bottleneck for many natural language processing (NLP) tasks.
-
1 Oct 2018 1 repository listedRecently introduced neural network parsers allow for new approaches to circumvent data sparsity issues by modeling character level information and by exploiting raw data in a semi-supervised setting.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections