{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/xtreme-a-massively-multilingual-multi-task-1","title":"XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation","arxiv_id":null,"date":"2020-01-01","proceeding":"ICML 2020 1","authors":["Junjie Hu","Sebastian Ruder","Aditya Siddhant","Graham Neubig","Orhan Firat","Melvin Johnson"],"abstract":"Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We will release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.\n\n","url_abs":"https://proceedings.icml.cc/static/paper_files/icml/2020/4220-Paper.pdf","url_pdf":"https://proceedings.icml.cc/static/paper_files/icml/2020/4220-Paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"xtreme-a-massively-multilingual-multi-task-1","repo_url":"https://github.com/google-research/xtreme","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"xtreme-a-massively-multilingual-multi-task-1","repo_url":"https://github.com/atreyasha/semantic-isometry-nmt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"cross-lingual-transfer","task_name":"Cross-Lingual Transfer"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-retrieval","task_name":"Sentence Retrieval"},{"task_slug":"zero-shot-cross-lingual-transfer","task_name":"Zero-Shot Cross-Lingual Transfer"}],"methods":[],"datasets_introduced":[{"slug":"xtreme","name":"XTREME","full_name":"Cross-Lingual Transfer Evaluation of Multilingual Encoders"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-cross-lingual-transfer-on-xtreme","task":"Zero-Shot Cross-Lingual Transfer","dataset":"XTREME","model":"mBERT","rank_in_archive_order":25,"of":25,"metrics":{"AVG":"59.6","Question Answering":"53.8","Sentence Retrieval":"47.7","Sentence-pair Classification":"73.7","Structured Prediction":"66.3"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}