{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/word-embeddings-via-tensor-factorization","title":"Word Embeddings via Tensor Factorization","arxiv_id":"1704.02686","date":"2017-04-10","proceeding":null,"authors":["Eric Bailey","Shuchin Aeron"],"abstract":"Most popular word embedding techniques involve implicit or explicit\nfactorization of a word co-occurrence based matrix into low rank factors. In\nthis paper, we aim to generalize this trend by using numerical methods to\nfactor higher-order word co-occurrence based arrays, or \\textit{tensors}. We\npresent four word embeddings using tensor factorization and analyze their\nadvantages and disadvantages. One of our main contributions is a novel joint\nsymmetric tensor factorization technique related to the idea of coupled tensor\nfactorization. We show that embeddings based on tensor factorization can be\nused to discern the various meanings of polysemous words without being\nexplicitly trained to do so, and motivate the intuition behind why this works\nin a way that doesn't with existing methods. We also modify an existing word\nembedding evaluation metric known as Outlier Detection [Camacho-Collados and\nNavigli, 2016] to evaluate the quality of the order-$N$ relations that a word\nembedding captures, and show that tensor-based methods outperform existing\nmatrix-based methods at this task. Experimentally, we show that all of our word\nembeddings either outperform or are competitive with state-of-the-art baselines\ncommonly used today on a variety of recent datasets. Suggested applications of\ntensor factorization-based word embeddings are given, and all source code and\npre-trained vectors are publicly available online.","url_abs":"http://arxiv.org/abs/1704.02686v2","url_pdf":"http://arxiv.org/pdf/1704.02686v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"word-embeddings-via-tensor-factorization","repo_url":"https://github.com/dnguyen1196/word-embedding-cp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"outlier-detection","task_name":"Outlier Detection"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}