{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-generative-word-embedding-model-and-its-low","title":"A Generative Word Embedding Model and its Low Rank Positive Semidefinite Solution","arxiv_id":"1508.03826","date":"2015-08-16","proceeding":"EMNLP 2015 9","authors":["Shaohua Li","Jun Zhu","Chunyan Miao"],"abstract":"Most existing word embedding methods can be categorized into Neural Embedding\nModels and Matrix Factorization (MF)-based methods. However some models are\nopaque to probabilistic interpretation, and MF-based methods, typically solved\nusing Singular Value Decomposition (SVD), may incur loss of corpus information.\nIn addition, it is desirable to incorporate global latent factors, such as\ntopics, sentiments or writing styles, into the word embedding model. Since\ngenerative models provide a principled way to incorporate latent factors, we\npropose a generative word embedding model, which is easy to interpret, and can\nserve as a basis of more sophisticated latent factor models. The model\ninference reduces to a low rank weighted positive semidefinite approximation\nproblem. Its optimization is approached by eigendecomposition on a submatrix,\nfollowed by online blockwise regression, which is scalable and avoids the\ninformation loss in SVD. In experiments on 7 common benchmark datasets, our\nvectors are competitive to word2vec, and better than other MF-based methods.","url_abs":"http://arxiv.org/abs/1508.03826v1","url_pdf":"http://arxiv.org/pdf/1508.03826v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-generative-word-embedding-model-and-its-low","repo_url":"https://github.com/askerlee/topicvec","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1508.03826","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}