{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-evaluation-of-neural-machine-translation","title":"An Evaluation of Neural Machine Translation Models on Historical Spelling Normalization","arxiv_id":"1806.05210","date":"2018-06-13","proceeding":"COLING 2018 8","authors":["Gongbo Tang","Fabienne Cap","Eva Pettersson","Joakim Nivre"],"abstract":"In this paper, we apply different NMT models to the problem of historical\nspelling normalization for five languages: English, German, Hungarian,\nIcelandic, and Swedish. The NMT models are at different levels, have different\nattention mechanisms, and different neural network architectures. Our results\nshow that NMT models are much better than SMT models in terms of character\nerror rate. The vanilla RNNs are competitive to GRUs/LSTMs in historical\nspelling normalization. Transformer models perform better only when provided\nwith more training data. We also find that subword-level models with a small\nsubword vocabulary are better than character-level models for low-resource\nlanguages. In addition, we propose a hybrid method which further improves the\nperformance of historical spelling normalization.","url_abs":"http://arxiv.org/abs/1806.05210v2","url_pdf":"http://arxiv.org/pdf/1806.05210v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-evaluation-of-neural-machine-translation","repo_url":"https://github.com/tanggongbo/normalization-NMT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"nmt","task_name":"NMT"},{"task_slug":"translation","task_name":"Translation"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.05210","atlas_url":"https://app.syntology.ai/?focus=1806.05210","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}