{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-unsupervised-and-customizable-misspelling","title":"An unsupervised and customizable misspelling generator for mining noisy health-related text sources","arxiv_id":"1806.00910","date":"2018-06-04","proceeding":null,"authors":["Abeed Sarker","Graciela Gonzalez-Hernandez"],"abstract":"In this paper, we present a customizable datacentric system that\nautomatically generates common misspellings for complex health-related terms.\nThe spelling variant generator relies on a dense vector model learned from\nlarge unlabeled text, which is used to find semantically close terms to the\noriginal/seed keyword, followed by the filtering of terms that are lexically\ndissimilar beyond a given threshold. The process is executed recursively,\nconverging when no new terms similar (lexically and semantically) to the seed\nkeyword are found. Weighting of intra-word character sequence similarities\nallows further problem-specific customization of the system. On a dataset\nprepared for this study, our system outperforms the current state-of-the-art\nfor medication name variant generation with best F1-score of 0.69 and\nF1/4-score of 0.78. Extrinsic evaluation of the system on a set of\ncancer-related terms showed an increase of over 67% in retrieval rate from\nTwitter posts when the generated variants are included. Our proposed spelling\nvariant generator has several advantages over the current state-of-the-art and\nother types of variant generators-(i) it is capable of filtering out lexically\nsimilar but semantically dissimilar terms, (ii) the number of variants\ngenerated is low as many low-frequency and ambiguous misspellings are filtered\nout, and (iii) the system is fully automatic, customizable and easily\nexecutable. While the base system is fully unsupervised, we show how\nsupervision maybe employed to adjust weights for task-specific customization.\nThe performance and significant relative simplicity of our proposed approach\nmakes it a much needed misspelling generation resource for health-related text\nmining from noisy sources. The source code for the system has been made\npublicly available for research purposes.","url_abs":"http://arxiv.org/abs/1806.00910v1","url_pdf":"http://arxiv.org/pdf/1806.00910v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-unsupervised-and-customizable-misspelling","repo_url":"https://bitbucket.org/asarker/qmisspell","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}