{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reproducing-and-learning-new-algebraic","title":"Reproducing and learning new algebraic operations on word embeddings using genetic programming","arxiv_id":"1702.05624","date":"2017-02-18","proceeding":null,"authors":["Roberto Santana"],"abstract":"Word-vector representations associate a high dimensional real-vector to every\nword from a corpus. Recently, neural-network based methods have been proposed\nfor learning this representation from large corpora. This type of\nword-to-vector embedding is able to keep, in the learned vector space, some of\nthe syntactic and semantic relationships present in the original word corpus.\nThis, in turn, serves to address different types of language classification\ntasks by doing algebraic operations defined on the vectors. The general\npractice is to assume that the semantic relationships between the words can be\ninferred by the application of a-priori specified algebraic operations. Our\ngeneral goal in this paper is to show that it is possible to learn methods for\nword composition in semantic spaces. Instead of expressing the compositional\nmethod as an algebraic operation, we will encode it as a program, which can be\nlinear, nonlinear, or involve more intricate expressions. More remarkably, this\nprogram will be evolved from a set of initial random programs by means of\ngenetic programming (GP). We show that our method is able to reproduce the same\nbehavior as human-designed algebraic operators. Using a word analogy task as\nbenchmark, we also show that GP-generated programs are able to obtain accuracy\nvalues above those produced by the commonly used human-designed rule for\nalgebraic manipulation of word vectors. Finally, we show the robustness of our\napproach by executing the evolved programs on the word2vec GoogleNews vectors,\nlearned over 3 billion running words, and assessing their accuracy in the same\nword analogy task.","url_abs":"http://arxiv.org/abs/1702.05624v1","url_pdf":"http://arxiv.org/pdf/1702.05624v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reproducing-and-learning-new-algebraic","repo_url":"https://github.com/rsantana-isg/GP_word2vec","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}