{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monoise-modeling-noise-using-a-modular","title":"MoNoise: Modeling Noise Using a Modular Normalization System","arxiv_id":"1710.03476","date":"2017-10-10","proceeding":null,"authors":["Rob van der Goot","Gertjan van Noord"],"abstract":"We propose MoNoise: a normalization model focused on generalizability and\nefficiency, it aims at being easily reusable and adaptable. Normalization is\nthe task of translating texts from a non- canonical domain to a more canonical\ndomain, in our case: from social media data to standard language. Our proposed\nmodel is based on a modular candidate generation in which each module is\nresponsible for a different type of normalization action. The most important\ngeneration modules are a spelling correction system and a word embeddings\nmodule. Depending on the definition of the normalization task, a static lookup\nlist can be crucial for performance. We train a random forest classifier to\nrank the candidates, which generalizes well to all different types of\nnormaliza- tion actions. Most features for the ranking originate from the\ngeneration modules; besides these features, N-gram features prove to be an\nimportant source of information. We show that MoNoise beats the\nstate-of-the-art on different normalization benchmarks for English and Dutch,\nwhich all define the task of normalization slightly different.","url_abs":"http://arxiv.org/abs/1710.03476v1","url_pdf":"http://arxiv.org/pdf/1710.03476v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"monoise-modeling-noise-using-a-modular","repo_url":"https://bitbucket.org/robvanderg/monoise","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"monoise-modeling-noise-using-a-modular","repo_url":"https://github.com/wesselreijngoud/masterthesis2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"lexical-normalization","task_name":"Lexical Normalization"},{"task_slug":"spelling-correction","task_name":"Spelling Correction"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/lexical-normalization-on-lexnorm","task":"Lexical Normalization","dataset":"LexNorm","model":"MoNoise","rank_in_archive_order":1,"of":4,"metrics":{"Accuracy":"87.63"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.03476","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}