{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-large-scale-comparison-of-historical-text","title":"A Large-Scale Comparison of Historical Text Normalization Systems","arxiv_id":"1904.02036","date":"2019-04-03","proceeding":"NAACL 2019 6","authors":["Marcel Bollmann"],"abstract":"There is no consensus on the state-of-the-art approach to historical text\nnormalization. Many techniques have been proposed, including rule-based\nmethods, distance metrics, character-based statistical machine translation, and\nneural encoder--decoder models, but studies have used different datasets,\ndifferent evaluation methods, and have come to different conclusions. This\npaper presents the largest study of historical text normalization done so far.\nWe critically survey the existing literature and report experiments on eight\nlanguages, comparing systems spanning all categories of proposed normalization\ntechniques, analysing the effect of training data quantity, and using different\nevaluation methods. The datasets and scripts are made publicly available.","url_abs":"http://arxiv.org/abs/1904.02036v1","url_pdf":"http://arxiv.org/pdf/1904.02036v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-large-scale-comparison-of-historical-text","repo_url":"https://github.com/coastalcph/histnorm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-large-scale-comparison-of-historical-text","repo_url":"https://github.com/SteffenEger/ocr_spelling_deuparl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"text-normalization","task_name":"Text Normalization"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.02036","atlas_url":"https://app.syntology.ai/?focus=1904.02036","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}