{"url":"/method/weight-tying","slug":"weight-tying","name":"Weight Tying","full_name":"Weight Tying","full_name_withheld":false,"description_markdown":"**Weight Tying** improves the performance of language models by tying (sharing) the weights of the embedding and [softmax](https://paperswithcode.com/method/softmax) layers. This method also massively reduces the total number of parameters in the language models that it is applied to. \r\n\r\nLanguage models are typically comprised of an embedding layer, followed by a number of [Transformer](https://paperswithcode.com/method/transformer) or [LSTM](https://paperswithcode.com/method/lstm) layers, which are finally followed by a softmax layer. Embedding layers learn word representations, such that similar words (in meaning) are represented by vectors that are near each other (in cosine distance). [Press & Wolf, 2016] showed that the softmax matrix, in which every word also has a vector representation, also exhibits this property. This leads them to propose to share the softmax and embedding matrices, which is done today in nearly all language models.  \r\n\r\nThis method was independently introduced by [Press & Wolf, 2016](https://paperswithcode.com/paper/using-the-output-embedding-to-improve) and [Inan et al, 2016](https://paperswithcode.com/paper/tying-word-vectors-and-word-classifiers-a).\r\n\r\nAdditionally, the Press & Wolf paper proposes Three-way Weight Tying, a method for NMT models in which the embedding matrix for the source language, the embedding matrix for the target language, and the softmax matrix for the target language are all tied. That method has been adopted by the Attention Is All You Need model and many other neural machine translation models.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Using the Output Embedding to Improve Language Models","paper":"/paper/using-the-output-embedding-to-improve","first_author":"Ofir Press","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/using-the-output-embedding-to-improve"},"source":{"url":"http://arxiv.org/abs/1608.05859v3","title":"Using the Output Embedding to Improve Language Models","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Parameter Sharing","url":"/methods/category/parameter-sharing","pwc_aliases":[]}],"n_papers_tagged":63,"archive_num_papers":63,"papers_newest_first":[{"paper":null,"title":"Advanced Deep Learning Techniques for Analyzing Earnings Call Transcripts: Methodologies and Applications","date":"2025-02-27","arxiv_id":"2503.01886","n_code_links":0,"syntology":null},{"paper":null,"title":"No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts","date":"2024-10-24","arxiv_id":"2410.19184","n_code_links":0,"syntology":null},{"paper":"/paper/rico-reddit-ideological-communities","title":"RICo: Reddit ideological communities","date":"2024-06-05","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/exploring-multi-level-threats-in-telegram","title":"Exploring Multi-Level Threats in Telegram Data with AI-Human Annotation: A Preliminary Study","date":"2023-12-15","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Illicit Darkweb Classification via Natural-language Processing: Classifying Illicit Content of Webpages based on Textual Information","date":"2023-12-08","arxiv_id":"2312.04944","n_code_links":0,"syntology":null},{"paper":null,"title":"Tied-Lora: Enhancing parameter efficiency of LoRA with weight tying","date":"2023-11-16","arxiv_id":"2311.09578","n_code_links":0,"syntology":null},{"paper":null,"title":"Headless Language Models: Learning without Predicting with Contrastive Weight Tying","date":"2023-09-15","arxiv_id":"2309.08351","n_code_links":0,"syntology":null},{"paper":"/paper/act3d-infinite-resolution-action-detection","title":"Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation","date":"2023-06-30","arxiv_id":"2306.17817","n_code_links":2,"syntology":null},{"paper":null,"title":"Explainable and High-Performance Hate and Offensive Speech Detection","date":"2022-06-26","arxiv_id":"2206.12983","n_code_links":0,"syntology":null},{"paper":"/paper/approximately-equivariant-networks-for","title":"Approximately Equivariant Networks for Imperfectly Symmetric Dynamics","date":"2022-01-28","arxiv_id":"2201.11969","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":"/paper/iiitt-dravidian-codemix-fire2021","title":"IIITT@Dravidian-CodeMix-FIRE2021: Transliterate or translate? Sentiment analysis of code-mixed text in Dravidian languages","date":"2021-11-15","arxiv_id":"2111.07906","n_code_links":1,"syntology":null},{"paper":"/paper/offensive-language-identification-in-low","title":"Offensive Language Identification in Low-resourced Code-mixed Dravidian languages using Pseudo-labeling","date":"2021-08-27","arxiv_id":"2108.12177","n_code_links":1,"syntology":null},{"paper":"/paper/sn-computer-science-towards-offensive","title":"Towards Offensive Language Identification for Tamil Code-Mixed YouTube Comments and Posts","date":"2021-08-24","arxiv_id":"2108.10939","n_code_links":1,"syntology":null},{"paper":null,"title":"Learning ULMFiT and Self-Distillation with Calibration for Medical Dialogue System","date":"2021-07-20","arxiv_id":"2107.09625","n_code_links":0,"syntology":null},{"paper":"/paper/whose-heritage-classification-of-unesco-world","title":"WHOSe Heritage: Classification of UNESCO World Heritage \"Outstanding Universal Value\" Documents with Soft Labels","date":"2021-04-12","arxiv_id":"2104.05547","n_code_links":1,"syntology":null},{"paper":"/paper/l3cubemahasent-a-marathi-tweet-based","title":"L3CubeMahaSent: A Marathi Tweet-based Sentiment Analysis Dataset","date":"2021-03-21","arxiv_id":"2103.11408","n_code_links":1,"syntology":null},{"paper":null,"title":"On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers","date":"2021-02-15","arxiv_id":"2102.07346","n_code_links":0,"syntology":null},{"paper":"/paper/indicnlp-kgp-at-dravidianlangtech-eacl2021","title":"indicnlp@kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages","date":"2021-02-14","arxiv_id":"2102.07150","n_code_links":1,"syntology":null},{"paper":"/paper/indicnlp-kgp-at-dravidianlangtech-eacl2021-1","title":"indicnlp@ kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages","date":"2021-02-14","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Train your classifier first: Cascade Neural Networks Training from upper layers to lower layers","date":"2021-02-09","arxiv_id":"2102.04697","n_code_links":0,"syntology":null},{"paper":null,"title":"Experimental Evaluation of Deep Learning models for Marathi Text Classification","date":"2021-01-13","arxiv_id":"2101.04899","n_code_links":0,"syntology":null},{"paper":"/paper/ladiff-ulmfit-a-layer-differentiated-training","title":"LaDiff ULMFiT: A Layer Differentiated training approach for ULMFiT","date":"2021-01-13","arxiv_id":"2101.04965","n_code_links":1,"syntology":null},{"paper":null,"title":"Post-Training Weighted Quantization of Neural Networks for Language Models","date":"2021-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/hinglishnlp-at-semeval-2020-task-9-fine-tuned","title":"HinglishNLP at SemEval-2020 Task 9: Fine-tuned Language Models for Hinglish Sentiment Detection","date":"2020-12-01","arxiv_id":null,"n_code_links":2,"syntology":null},{"paper":null,"title":"Smash at SemEval-2020 Task 7: Optimizing the Hyperparameters of ERNIE 2.0 for Humor Ranking and Rating","date":"2020-12-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Palomino-Ochoa at SemEval-2020 Task 9: Robust System based on Transformer for Code-Mixed Sentiment Classification","date":"2020-11-18","arxiv_id":"2011.09448","n_code_links":0,"syntology":null},{"paper":"/paper/pagsusuri-ng-rnn-based-transfer-learning","title":"Pagsusuri ng RNN-based Transfer Learning Technique sa Low-Resource Language","date":"2020-10-13","arxiv_id":"2010.06447","n_code_links":2,"syntology":null},{"paper":null,"title":"Gauravarora@HASOC-Dravidian-CodeMix-FIRE2020: Pre-training ULMFiT on Synthetically Generated Code-Mixed Data for Hate Speech Detection","date":"2020-10-05","arxiv_id":"2010.02094","n_code_links":0,"syntology":null},{"paper":null,"title":"Fine-tuning Pre-trained Contextual Embeddings for Citation Content Analysis in Scholarly Publication","date":"2020-09-12","arxiv_id":"2009.05836","n_code_links":0,"syntology":null},{"paper":"/paper/hinglishnlp-fine-tuned-language-models-for","title":"HinglishNLP: Fine-tuned Language Models for Hinglish Sentiment Detection","date":"2020-08-22","arxiv_id":"2008.09820","n_code_links":2,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":23},{"task":"/task/language-modeling","name":"Language Modeling","papers":21},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":16},{"task":"/task/classification","name":"General Classification","papers":15},{"task":"/task/text-classification","name":"Text Classification","papers":15},{"task":"/task/text-classification-1","name":"text-classification","papers":11},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":9},{"task":"/task/classification-1","name":"Classification","papers":8},{"task":"/task/translation","name":"Translation","papers":7},{"task":"/task/machine-translation","name":"Machine Translation","papers":5},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":5},{"task":"/task/language-identification","name":"Language Identification","papers":4},{"task":"/task/decision-making","name":"Decision Making","papers":3},{"task":"/task/hate-speech-detection","name":"Hate Speech Detection","papers":3},{"task":"/task/sentence","name":"Sentence","papers":3},{"task":"/task/sentiment-classification","name":"Sentiment Classification","papers":3},{"task":"/task/articles","name":"Articles","papers":2},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":2},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/image-classification","name":"Image Classification","papers":2}],"tasks_shown":20,"n_tasks":76,"usage_by_year":[{"year":"2016","papers":2},{"year":"2017","papers":3},{"year":"2018","papers":4},{"year":"2019","papers":16},{"year":"2020","papers":15},{"year":"2021","papers":13},{"year":"2022","papers":2},{"year":"2023","papers":5},{"year":"2024","papers":2},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/weight-tying"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}