Papers › InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training

InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training

15 Jul 2020NAACL 2021 4arXiv:2007.07834archive 2025-07-28

Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, He-Yan Huang, Ming Zhou

In this work, we present an information-theoretic framework that formulates cross-lingual language model pre-training as maximizing mutual information between multilingual-multi-granularity texts. The unified view helps us to better understand the existing methods for learning cross-lingual representations. More importantly, inspired by the framework, we propose a new pre-training task based on contrastive learning. Specifically, we regard a bilingual sentence pair as two views of the same meaning and encourage their encoded representations to be more similar than the negative examples. By leveraging both monolingual and parallel corpora, we jointly train the pretext tasks to improve the cross-lingual transferability of pre-trained models. Experimental results on several benchmarks show that our approach achieves considerably better performance. The code and pre-trained models are available at https://aka.ms/infoxlm.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

CZWin32768/xnlg mentioned on GitHubpytorch report
facebookresearch/data2vec_vision mentioned on GitHubpytorch report
jiamingkong/infoxlm_paddle mentioned on GitHubpaddle report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningCross-Lingual TransferLanguage ModelingLanguage ModellingSentence

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Cross-Lingual Transfer XTREME T-ULRv2 + StableTune Avg 80.7 #16 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME T-ULRv2 + StableTune Question Answering 72.9 #16 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME T-ULRv2 + StableTune Sentence Retrieval 89.3 #16 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME T-ULRv2 + StableTune Sentence-pair Classification 88.8 #16 of 25 Archive leaderboard report
Zero-Shot Cross-Lingual Transfer XTREME T-ULRv2 + StableTune Structured Prediction 75.4 #16 of 25 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEContrastive Multiview CodingDense ConnectionsDropoutInfoNCELayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections