{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adapting-word-embeddings-to-new-languages","title":"Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations","arxiv_id":"1808.09500","date":"2018-08-28","proceeding":"EMNLP 2018 10","authors":["Aditi Chaudhary","Chunting Zhou","Lori Levin","Graham Neubig","David R. Mortensen","Jaime G. Carbonell"],"abstract":"Much work in Natural Language Processing (NLP) has been for resource-rich\nlanguages, making generalization to new, less-resourced languages challenging.\nWe present two approaches for improving generalization to low-resourced\nlanguages by adapting continuous word representations using linguistically\nmotivated subword units: phonemes, morphemes and graphemes. Our method requires\nneither parallel corpora nor bilingual dictionaries and provides a significant\ngain in performance over previous methods relying on these resources. We\ndemonstrate the effectiveness of our approaches on Named Entity Recognition for\nfour languages, namely Uyghur, Turkish, Bengali and Hindi, of which Uyghur and\nBengali are low resource languages, and also perform experiments on Machine\nTranslation. Exploiting subwords with transfer learning gives us a boost of\n+15.2 NER F1 for Uyghur and +9.7 F1 for Bengali. We also show improvements in\nthe monolingual setting where we achieve (avg.) +3 F1 and (avg.) +1.35 BLEU.","url_abs":"http://arxiv.org/abs/1808.09500v1","url_pdf":"http://arxiv.org/pdf/1808.09500v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adapting-word-embeddings-to-new-languages","repo_url":"https://github.com/Aditi138/Embeddings","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":null,"task_name":"Avg"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"cg","task_name":"NER"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1808.09500","atlas_url":"https://app.syntology.ai/?focus=1808.09500","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}