{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-character-level-approach-to-the-text","title":"A Character-Level Approach to the Text Normalization Problem Based on a New Causal Encoder","arxiv_id":"1903.02642","date":"2019-03-06","proceeding":null,"authors":["Adrián Javaloy Bornás","Ginés García Mateos"],"abstract":"Text normalization is a ubiquitous process that appears as the first step of\nmany Natural Language Processing problems. However, previous Deep Learning\napproaches have suffered from so-called silly errors, which are undetectable on\nunsupervised frameworks, making those models unsuitable for deployment. In this\nwork, we make use of an attention-based encoder-decoder architecture that\novercomes these undetectable errors by using a fine-grained character-level\napproach rather than a word-level one. Furthermore, our new general-purpose\nencoder based on causal convolutions, called Causal Feature Extractor (CFE), is\nintroduced and compared to other common encoders. The experimental results show\nthe feasibility of this encoder, which leverages the attention mechanisms the\nmost and obtains better results in terms of accuracy, number of parameters and\nconvergence time. While our method results in a slightly worse initial accuracy\n(92.74%), errors can be automatically detected and, thus, more readily solved,\nobtaining a more robust model for deployment. Furthermore, there is still\nplenty of room for future improvements that will push even further these\nadvantages.","url_abs":"http://arxiv.org/abs/1903.02642v1","url_pdf":"http://arxiv.org/pdf/1903.02642v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-character-level-approach-to-the-text","repo_url":"https://github.com/adrianjav/text-normalization-preprocess","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"text-normalization","task_name":"Text Normalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}