{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/machine-translation-of-low-resource-spoken","title":"Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German","arxiv_id":"1710.11035","date":"2017-10-30","proceeding":"LREC 2018 5","authors":["Pierre-Edouard Honnet","Andrei Popescu-Belis","Claudiu Musat","Michael Baeriswyl"],"abstract":"The goal of this work is to design a machine translation (MT) system for a\nlow-resource family of dialects, collectively known as Swiss German, which are\nwidely spoken in Switzerland but seldom written. We collected a significant\nnumber of parallel written resources to start with, up to a total of about 60k\nwords. Moreover, we identified several other promising data sources for Swiss\nGerman. Then, we designed and compared three strategies for normalizing Swiss\nGerman input in order to address the regional diversity. We found that\ncharacter-based neural MT was the best solution for text normalization. In\ncombination with phrase-based statistical MT, our solution reached 36% BLEU\nscore when translating from the Bernese dialect. This value, however, decreases\nas the testing data becomes more remote from the training one, geographically\nand topically. These resources and normalization techniques are a first step\ntowards full MT of Swiss German dialects.","url_abs":"http://arxiv.org/abs/1710.11035v2","url_pdf":"http://arxiv.org/pdf/1710.11035v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"machine-translation-of-low-resource-spoken","repo_url":"https://github.com/Kyubyong/quasi-rnn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"text-normalization","task_name":"Text Normalization"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}