{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/characterizing-the-impact-of-geometric","title":"Characterizing the impact of geometric properties of word embeddings on task performance","arxiv_id":"1904.04866","date":"2019-04-09","proceeding":"WS 2019 6","authors":["Brendan Whitaker","Denis Newman-Griffis","Aparajita Haldar","Hakan Ferhatosmanoglu","Eric Fosler-Lussier"],"abstract":"Analysis of word embedding properties to inform their use in downstream NLP\ntasks has largely been studied by assessing nearest neighbors. However,\ngeometric properties of the continuous feature space contribute directly to the\nuse of embedding features in downstream models, and are largely unexplored. We\nconsider four properties of word embedding geometry, namely: position relative\nto the origin, distribution of features in the vector space, global pairwise\ndistances, and local pairwise distances. We define a sequence of\ntransformations to generate new embeddings that expose subsets of these\nproperties to downstream models and evaluate change in task performance to\nunderstand the contribution of each property to NLP models. We transform\npublicly available pretrained embeddings from three popular toolkits (word2vec,\nGloVe, and FastText) and evaluate on a variety of intrinsic tasks, which model\nlinguistic information in the vector space, and extrinsic tasks, which use\nvectors as input to machine learning models. We find that intrinsic evaluations\nare highly sensitive to absolute position, while extrinsic tasks rely primarily\non local similarity. Our findings suggest that future embedding models and\npost-processing techniques should focus primarily on similarity to nearby\npoints in vector space.","url_abs":"http://arxiv.org/abs/1904.04866v1","url_pdf":"http://arxiv.org/pdf/1904.04866v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"characterizing-the-impact-of-geometric","repo_url":"https://github.com/OSU-slatelab/geometric-embedding-properties","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"characterizing-the-impact-of-geometric","repo_url":"https://github.com/drgriffis/Extrinsic-Evaluation-tasks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":null,"task_name":"Position"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}