{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transfer-learning-in-multilingual-neural","title":"Transfer Learning in Multilingual Neural Machine Translation with Dynamic Vocabulary","arxiv_id":"1811.01137","date":"2018-11-03","proceeding":"IWSLT (EMNLP) 2018 10","authors":["Surafel M. Lakew","Aliia Erofeeva","Matteo Negri","Marcello Federico","Marco Turchi"],"abstract":"We propose a method to transfer knowledge across neural machine translation\n(NMT) models by means of a shared dynamic vocabulary. Our approach allows to\nextend an initial model for a given language pair to cover new languages by\nadapting its vocabulary as long as new data become available (i.e., introducing\nnew vocabulary items if they are not included in the initial model). The\nparameter transfer mechanism is evaluated in two scenarios: i) to adapt a\ntrained single language NMT system to work with a new language pair and ii) to\ncontinuously add new language pairs to grow to a multilingual NMT system. In\nboth the scenarios our goal is to improve the translation performance, while\nminimizing the training convergence time. Preliminary experiments spanning five\nlanguages with different training data sizes (i.e., 5k and 50k parallel\nsentences) show a significant performance gain ranging from +3.85 up to +13.63\nBLEU in different language directions. Moreover, when compared with training an\nNMT model from scratch, our transfer-learning approach allows us to reach\nhigher performance after training up to 4% of the total training steps.","url_abs":"http://arxiv.org/abs/1811.01137v1","url_pdf":"http://arxiv.org/pdf/1811.01137v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transfer-learning-in-multilingual-neural","repo_url":"https://github.com/surafelml/Afro-NMT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"transfer-learning-in-multilingual-neural","repo_url":"https://github.com/surafelml/adapt-mnmt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"nmt","task_name":"NMT"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.01137","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}