Papers › Learning Word Vectors for 157 Languages

Learning Word Vectors for 157 Languages

19 Feb 2018LREC 2018 5arXiv:1802.06893archive 2025-07-28

Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, Tomas Mikolov

Distributed word representations, or word vectors, have recently been applied to many tasks in natural language processing, leading to state-of-the-art performance. A key ingredient to the successful application of these representations is to train them on very large corpora, and use these pre-trained models in downstream tasks. In this paper, we describe how we trained such high quality word representations for 157 languages. We used two sources of data to train these models: the free online encyclopedia Wikipedia and data from the common crawl project. We also introduce three new word analogy datasets to evaluate these word vectors, for French, Hindi and Polish. Finally, we evaluate our pre-trained word vectors on 10 languages for which evaluation datasets exists, showing very strong performance compared to previous models.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

KMicha/MachineLearning mentioned on GitHubApache-2.0 report
dzieciou/lemmatizer-pl mentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Only Connect Walls Dataset Task 1 (Grouping)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (Crawl) Wasserstein Distance (WD) 84.2 ± .5 #12 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (Crawl) # Correct Groups 80 ± 4 #12 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (Crawl) # Solved Walls 0 ± 0 #12 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (Crawl) Adjusted Mutual Information (AMI) 18.4 ± .4 #12 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (Crawl) Adjusted Rand Index (ARI) 15.2 ± .3 #12 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (Crawl) Fowlkes Mallows Score (FMS) 32.1 ± .3 #12 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (News) Wasserstein Distance (WD) 85.5 ± .5 #15 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (News) # Correct Groups 62 ± 3 #15 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (News) # Solved Walls 0 ± 0 #15 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (News) Adjusted Mutual Information (AMI) 15.8 ± .3 #15 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (News) Adjusted Rand Index (ARI) 13.0 ± .2 #15 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW FastText (News) Fowlkes Mallows Score (FMS) 30.4 ± .2 #15 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections