Papers › Unsupervised Cross-lingual Representation Learning at Scale
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov
This paper shows that pretraining multilingual language models at scale leads to significant performance gains for a wide range of cross-lingual transfer tasks. We train a Transformer-based masked language model on one hundred languages, using more than two terabytes of filtered CommonCrawl data. Our model, dubbed XLM-R, significantly outperforms multilingual BERT (mBERT) on a variety of cross-lingual benchmarks, including +14.6% average accuracy on XNLI, +13% average F1 score on MLQA, and +2.4% F1 score on NER. XLM-R performs particularly well on low-resource languages, improving 15.7% in XNLI accuracy for Swahili and 11.4% for Urdu over previous XLM models. We also present a detailed empirical analysis of the key factors that are required to achieve these gains, including the trade-offs between (1) positive transfer and capacity dilution and (2) the performance of high and low resource languages at scale. Finally, we show, for the first time, the possibility of multilingual modeling without sacrificing per-language performance; XLM-R is very competitive with strong monolingual models on the GLUE and XNLI benchmarks. We will make our code, data and models publicly available.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="1911.02116")
Code
Syntology Ran 25 of 59 code samples harvested from 9 repositories linked to this paper; 34 have no recorded run. Of those that ran: 4 ran · honoured contract; 1 ran · violated contract; 2 ran · our draft was wrong; 3 ran · fixture could not drive it; 15 ran with no contract checked.
By repository: official repository: 15 samples from 1 repository, 0 ran; community (archive-listed): 40 samples from 8 repositories, 21 ran; 4 identical to code first harvested elsewhere. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
35 repositories listed; official and paper-mentioned ones first.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
59 samples harvested; 25 ran; 4 honoured the contract we drafted; 34 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 52 of the 59 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 9 repositories linked to this paper, official or community; each sample names its own and says which. Some samples are identical code Syntology first harvested from another repository; for those, this paper's copy is not located and its licence is not recorded. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
b9f2456deb48efe3 · report
9a3ea499eb0114dc · report
c8b6d34869aa3762 · report
4bb30f62c8aa7df1 · report
6cb4b3cddb2b34b6 · report
78ae9912884277ad · report
3143d897968b04d0 · report
1e7f7cbd3125f871 · report
1f765197d3df9584 · report
aeaad31adb65d876 · report
50c73a5e3a282bc8 · report
0d75948d09185811 · report
8e7895f398977866 · report
cda9917a1ba9c984 · report
8e1af02efb15e082 · report
ea265a8f7f9d376b · report
2d39718257c802dc · report
d470784790e9dc66 · report
4a9d2d7dd384f8bf · report
59c6324aa51d05e7 · report
87278bd6d018466e · report
8ed44ce7206b1f24 · report
d204f4ff6c43fc6d · report
ba5fdac8a37a5455 · report
b8082687538a83e1 · report
ddc3ebfb70be4b09 · report
8ba7ef0726990c84 · report
2c9fdf8b95e37258 · report
72cc8733f4306eba · report
38525832e182b5f3 · report
424066fae6f1c350 · report
e7bd7636a1df749c · report
1d77d0309d035999 · report
91283fd5a725ec99 · report
3fa144c23db44a7e · report
a05791210b07b83d · report
bee8cacb03b2b055 · report
33c46f658371888d · report
ce7a4a01c9b960a2 · report
d30d1ecf1454760e · report
3d693a4a286601f2 · report
c4f0cfb3b4341157 · report
f8da1eb21868463a · report
200a1dfd844a9fe4 · report
e836ec1da3a517b3 · report
68f606f965d049f0 · report
292808e3a5d973ed · report
12b4ab9bb368a1e8 · report
99c39b47ca18e1b5 · report
1e9b7714930dc642 · report
b2b68b9410ef8d8b · report
80bf3f1cdd634770 · report
81d3a619bdd21b88 · report
ded0b9ac1f3c803c · report
41f2f12e6d5e6d33 · report
b5254f10aa821ab8 · report
cc66d68d2046d6a3 · report
66c21a3911091d29 · report
b2470f4f0c09db41 · report
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
No leaderboard rows for this paper in the archive.
Methods
Introduced by this paper: XLM-R
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections