{"url":"/task/lexical-normalization","name":"Lexical Normalization","slug":"lexical-normalization","description_markdown":"Lexical normalization is the task of translating/transforming a non standard text to a standard register.\n\nExample:\n\n```\nnew pix comming tomoroe\nnew pictures coming tomorrow\n```\n\nDatasets usually consists of tweets, since these naturally contain a fair amount of \nthese phenomena.\n\nFor lexical normalization, only replacements on the word-level are annotated.\nSome corpora include annotation for 1-N and N-1 replacements. However, word\ninsertion/deletion and reordering is not part of the task.","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":47,"papers_with_code":16,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":1,"subtasks":1,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/lexical-normalization-on-lexnorm","slug":"lexical-normalization-on-lexnorm","dataset":"LexNorm","dataset_url":null,"rows_in_archive":4,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"MoNoise","paper_title":"MoNoise: Modeling Noise Using a Modular Normalization System","paper_url":"/paper/monoise-modeling-noise-using-a-modular","paper_date":"2017-10-10","arxiv_id":"1710.03476","code_links":[{"title":"wesselreijngoud/masterthesis2019","url":"https://github.com/wesselreijngoud/masterthesis2019"},{"title":"robvanderg/monoise","url":"https://bitbucket.org/robvanderg/monoise"}],"syntology":null}}],"datasets":[{"url":"/dataset/multisenti","name":"MultiSenti","full_name":"","num_papers_in_archive":1}],"subtasks":[{"url":"/task/pronunciation-dictionary-creation","name":"Pronunciation Dictionary Creation"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":16,"of":16,"tagged_in_all":47,"items":[{"url":"/paper/monoise-modeling-noise-using-a-modular","title":"MoNoise: Modeling Noise Using a Modular Normalization System","date":"2017-10-10","arxiv_id":"1710.03476","repositories_listed":2,"syntology":null},{"url":"/paper/visolex-an-open-source-repository-for","title":"ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization","date":"2025-01-13","arxiv_id":"2501.07020","repositories_listed":1,"syntology":null},{"url":"/paper/vilexnorm-a-lexical-normalization-corpus-for","title":"ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text","date":"2024-01-29","arxiv_id":"2401.16403","repositories_listed":1,"syntology":null},{"url":"/paper/automatic-textual-normalization-for-hate","title":"Automatic Textual Normalization for Hate Speech Detection","date":"2023-11-12","arxiv_id":"2311.06851","repositories_listed":1,"syntology":null},{"url":"/paper/increasing-robustness-for-cross-domain","title":"Increasing Robustness for Cross-domain Dialogue Act Classification on Social Media Data","date":"2022-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/multilexnorm-a-shared-task-on-multilingual","title":"MultiLexNorm: A Shared Task on Multilingual Lexical Normalization","date":"2021-11-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/ufal-at-multilexnorm-2021-improving","title":"ÚFAL at MultiLexNorm 2021: Improving Multilingual Lexical Normalization by Fine-tuning ByT5","date":"2021-10-28","arxiv_id":"2110.15248","repositories_listed":1,"syntology":null},{"url":"/paper/dan-danish-nested-named-entities-and-lexical","title":"DaN+: Danish Nested Named Entities and Lexical Normalization","date":"2021-05-24","arxiv_id":"2105.11301","repositories_listed":1,"syntology":null},{"url":"/paper/user-generated-text-corpus-for-evaluating","title":"User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization","date":"2021-04-08","arxiv_id":"2104.03523","repositories_listed":1,"syntology":null},{"url":"/paper/lexical-normalization-for-code-switched-data-1","title":"Lexical Normalization for Code-switched Data and its Effect on POS Tagging","date":"2021-04-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/a-clustering-framework-for-lexical","title":"A Clustering Framework for Lexical Normalization of Roman Urdu","date":"2020-03-31","arxiv_id":"2004.00088","repositories_listed":1,"syntology":null},{"url":"/paper/adapting-deep-learning-for-sentiment","title":"Adapting Deep Learning for Sentiment Classification of Code-Switched Informal Short Text","date":"2020-01-04","arxiv_id":"2001.01047","repositories_listed":1,"syntology":null},{"url":"/paper/a-multi-cascaded-deep-model-for-bilingual-sms","title":"A Multi-cascaded Deep Model for Bilingual SMS Classification","date":"2019-11-29","arxiv_id":"1911.13066","repositories_listed":1,"syntology":null},{"url":"/paper/monoise-a-multi-lingual-and-easy-to-use","title":"MoNoise: A Multi-lingual and Easy-to-use Lexical Normalization Tool","date":"2019-07-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/adapting-sequence-to-sequence-models-for-text","title":"Adapting Sequence to Sequence models for Text Normalization in Social Media","date":"2019-04-12","arxiv_id":"1904.06100","repositories_listed":1,"syntology":null},{"url":"/paper/modeling-input-uncertainty-in-neural-network","title":"Modeling Input Uncertainty in Neural Network Dependency Parsing","date":"2018-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null}],"syntology_records":0,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}