{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-word-vectors-for-157-languages","title":"Learning Word Vectors for 157 Languages","arxiv_id":"1802.06893","date":"2018-02-19","proceeding":"LREC 2018 5","authors":["Edouard Grave","Piotr Bojanowski","Prakhar Gupta","Armand Joulin","Tomas Mikolov"],"abstract":"Distributed word representations, or word vectors, have recently been applied\nto many tasks in natural language processing, leading to state-of-the-art\nperformance. A key ingredient to the successful application of these\nrepresentations is to train them on very large corpora, and use these\npre-trained models in downstream tasks. In this paper, we describe how we\ntrained such high quality word representations for 157 languages. We used two\nsources of data to train these models: the free online encyclopedia Wikipedia\nand data from the common crawl project. We also introduce three new word\nanalogy datasets to evaluate these word vectors, for French, Hindi and Polish.\nFinally, we evaluate our pre-trained word vectors on 10 languages for which\nevaluation datasets exists, showing very strong performance compared to\nprevious models.","url_abs":"http://arxiv.org/abs/1802.06893v2","url_pdf":"http://arxiv.org/pdf/1802.06893v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-word-vectors-for-157-languages","repo_url":"https://github.com/KMicha/MachineLearning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"learning-word-vectors-for-157-languages","repo_url":"https://github.com/dzieciou/lemmatizer-pl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"task-1-grouping","task_name":"Only Connect Walls Dataset Task 1 (Grouping)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/task-1-grouping-on-ocw","task":"Only Connect Walls Dataset Task 1 (Grouping)","dataset":"OCW","model":"FastText (Crawl)","rank_in_archive_order":12,"of":22,"metrics":{" Wasserstein Distance (WD)":"84.2 ± .5","# Correct Groups":"80 ± 4","# Solved Walls":"0 ± 0","Adjusted Mutual Information (AMI)":"18.4 ± .4","Adjusted Rand Index (ARI)":"15.2 ± .3","Fowlkes Mallows Score (FMS)":"32.1 ± .3"},"uses_additional_data":true},{"leaderboard":"/sota/task-1-grouping-on-ocw","task":"Only Connect Walls Dataset Task 1 (Grouping)","dataset":"OCW","model":"FastText (News)","rank_in_archive_order":15,"of":22,"metrics":{" Wasserstein Distance (WD)":"85.5 ± .5","# Correct Groups":"62 ± 3","# Solved Walls":"0 ± 0","Adjusted Mutual Information (AMI)":" 15.8 ± .3","Adjusted Rand Index (ARI)":"13.0 ± .2","Fowlkes Mallows Score (FMS)":"30.4 ± .2"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/1802.06893","atlas_url":"https://app.syntology.ai/?focus=1802.06893","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}