{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-to-properly-evaluate-cross-lingual-word","title":"How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions","arxiv_id":"1902.00508","date":"2019-02-01","proceeding":"ACL 2019 7","authors":["Goran Glavas","Robert Litschko","Sebastian Ruder","Ivan Vulic"],"abstract":"Cross-lingual word embeddings (CLEs) enable multilingual modeling of meaning\nand facilitate cross-lingual transfer of NLP models. Despite their ubiquitous\nusage in downstream tasks, recent increasingly popular projection-based CLE\nmodels are almost exclusively evaluated on a single task only: bilingual\nlexicon induction (BLI). Even BLI evaluations vary greatly, hindering our\nability to correctly interpret performance and properties of different CLE\nmodels. In this work, we make the first step towards a comprehensive evaluation\nof cross-lingual word embeddings. We thoroughly evaluate both supervised and\nunsupervised CLE models on a large number of language pairs in the BLI task and\nthree downstream tasks, providing new insights concerning the ability of\ncutting-edge CLE models to support cross-lingual NLP. We empirically\ndemonstrate that the performance of CLE models largely depends on the task at\nhand and that optimizing CLE models for BLI can result in deteriorated\ndownstream performance. We indicate the most robust supervised and unsupervised\nCLE models and emphasize the need to reassess existing baselines, which still\ndisplay competitive performance across the board. We hope that our work will\ncatalyze further work on CLE evaluation and model analysis.","url_abs":"http://arxiv.org/abs/1902.00508v1","url_pdf":"http://arxiv.org/pdf/1902.00508v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-to-properly-evaluate-cross-lingual-word","repo_url":"https://github.com/codogogo/xling-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"bilingual-lexicon-induction","task_name":"Bilingual Lexicon Induction"},{"task_slug":"cross-lingual-transfer","task_name":"Cross-Lingual Transfer"},{"task_slug":"cross-lingual-word-embeddings","task_name":"Cross-Lingual Word Embeddings"},{"task_slug":"misconceptions","task_name":"Misconceptions"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1902.00508","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}