{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-k-way-d-dimensional-discrete-codes","title":"Learning K-way D-dimensional Discrete Codes for Compact Embedding Representations","arxiv_id":"1806.09464","date":"2018-06-21","proceeding":"ICML 2018 7","authors":["Ting Chen","Martin Renqiang Min","Yizhou Sun"],"abstract":"Conventional embedding methods directly associate each symbol with a\ncontinuous embedding vector, which is equivalent to applying a linear\ntransformation based on a \"one-hot\" encoding of the discrete symbols. Despite\nits simplicity, such approach yields the number of parameters that grows\nlinearly with the vocabulary size and can lead to overfitting. In this work, we\npropose a much more compact K-way D-dimensional discrete encoding scheme to\nreplace the \"one-hot\" encoding. In the proposed \"KD encoding\", each symbol is\nrepresented by a $D$-dimensional code with a cardinality of $K$, and the final\nsymbol embedding vector is generated by composing the code embedding vectors.\nTo end-to-end learn semantically meaningful codes, we derive a relaxed discrete\noptimization approach based on stochastic gradient descent, which can be\ngenerally applied to any differentiable computational graph with an embedding\nlayer. In our experiments with various applications from natural language\nprocessing to graph convolutional networks, the total size of the embedding\nlayer can be reduced up to 98\\% while achieving similar or better performance.","url_abs":"http://arxiv.org/abs/1806.09464v1","url_pdf":"http://arxiv.org/pdf/1806.09464v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-k-way-d-dimensional-discrete-codes","repo_url":"https://github.com/chentingpc/kdcode-lm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.09464","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}