Papers › XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual...

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation

1 Jan 2020ICML 2020 1archive 2025-07-28

Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, Melvin Johnson

Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We will release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.

PaperPDFConference PDFCode

Code

google-research/xtreme officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Lingual TransferRetrievalSentenceSentence RetrievalZero-Shot Cross-Lingual Transfer

Datasets

Introduced by this paper, per the archive.

XTREME

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Cross-Lingual Transfer XTREME mBERT AVG 59.6 #25 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME mBERT Question Answering 53.8 #25 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME mBERT Sentence Retrieval 47.7 #25 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME mBERT Sentence-pair Classification 73.7 #25 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME mBERT Structured Prediction 66.3 #25 of 25 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections