{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-to-evaluate-word-embeddings-on-importance","title":"How to evaluate word embeddings? On importance of data efficiency and simple supervised tasks","arxiv_id":"1702.02170","date":"2017-02-07","proceeding":null,"authors":["Stanisław Jastrzebski","Damian Leśniak","Wojciech Marian Czarnecki"],"abstract":"Maybe the single most important goal of representation learning is making\nsubsequent learning faster. Surprisingly, this fact is not well reflected in\nthe way embeddings are evaluated. In addition, recent practice in word\nembeddings points towards importance of learning specialized representations.\nWe argue that focus of word representation evaluation should reflect those\ntrends and shift towards evaluating what useful information is easily\naccessible. Specifically, we propose that evaluation should focus on data\nefficiency and simple supervised tasks, where the amount of available data is\nvaried and scores of a supervised model are reported for each subset (as\ncommonly done in transfer learning).\n  In order to illustrate significance of such analysis, a comprehensive\nevaluation of selected word embeddings is presented. Proposed approach yields a\nmore complete picture and brings new insight into performance characteristics,\nfor instance information about word similarity or analogy tends to be\nnon--linearly encoded in the embedding space, which questions the cosine-based,\nunsupervised, evaluation methods. All results and analysis scripts are\navailable online.","url_abs":"http://arxiv.org/abs/1702.02170v1","url_pdf":"http://arxiv.org/pdf/1702.02170v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-to-evaluate-word-embeddings-on-importance","repo_url":"https://github.com/NilsRethmeier/MoRTy","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"how-to-evaluate-word-embeddings-on-importance","repo_url":"https://github.com/PyENE/meta-word-embedding","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"how-to-evaluate-word-embeddings-on-importance","repo_url":"https://github.com/avsilva/sparse-nlp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"how-to-evaluate-word-embeddings-on-importance","repo_url":"https://github.com/kudkudak/word-embeddings-benchmarks","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"word-similarity","task_name":"Word Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1702.02170","atlas_url":"https://app.syntology.ai/?focus=1702.02170","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}