{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/beyond-weight-tying-learning-joint-input","title":"Beyond Weight Tying: Learning Joint Input-Output Embeddings for Neural Machine Translation","arxiv_id":"1808.10681","date":"2018-08-31","proceeding":"WS 2018 10","authors":["Nikolaos Pappas","Lesly Miculicich Werlen","James Henderson"],"abstract":"Tying the weights of the target word embeddings with the target word\nclassifiers of neural machine translation models leads to faster training and\noften to better translation quality. Given the success of this parameter\nsharing, we investigate other forms of sharing in between no sharing and hard\nequality of parameters. In particular, we propose a structure-aware output\nlayer which captures the semantic structure of the output space of words within\na joint input-output embedding. The model is a generalized form of weight tying\nwhich shares parameters but allows learning a more flexible relationship with\ninput word embeddings and allows the effective capacity of the output layer to\nbe controlled. In addition, the model shares weights across output classifiers\nand translation contexts which allows it to better leverage prior knowledge\nabout them. Our evaluation on English-to-Finnish and English-to-German datasets\nshows the effectiveness of the method against strong encoder-decoder baselines\ntrained with or without weight tying.","url_abs":"http://arxiv.org/abs/1808.10681v1","url_pdf":"http://arxiv.org/pdf/1808.10681v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"beyond-weight-tying-learning-joint-input","repo_url":"https://github.com/idiap/joint-embedding-nmt","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.10681","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}