Browse State-of-the-Art › Cross-Lingual Transfer
Cross-Lingual Transfer
333 papers with code · 1 benchmark · 15 datasets archive 2025-07-28
Cross-lingual transfer refers to transfer learning using data and models available for one language for which ample such resources are available (e.g., English) to solve tasks in another, commonly more low-resource, language.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| XCOPA (6 rows) | PaLM 2 (few-shot) | PaLM 2 Technical Report | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
15 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 333 papers with code (782 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Nov 2019 35 repositories listed Syntology ran 25 of 59 samples · 34 unverified · 52 pointer-only (licence)We also present a detailed empirical analysis of the key factors that are required to achieve these gains, including the trade-offs between (1) positive transfer and capacity dilution and (2) the performance of high and…
-
26 Dec 2018 13 repositories listed Syntology ran 4 of 10 samples · 6 unverified · 4 pointer-only (licence)We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts.
-
20 Apr 2023 6 repositories listedThis can for instance be observed when finetuning PLMs on one language and evaluating them on data in a closely related language variety with no standardized orthography.
-
16 Dec 2021 6 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedIn this work, we explore the limits of contrastive learning as a way to train unsupervised dense retrievers and show that it leads to strong performance in various retrieval settings.
-
23 Oct 2022 4 repositories listedZero-resource cross-lingual transfer approaches aim to apply supervised models from a source language to unlabelled target languages.
-
15 Oct 2021 4 repositories listedWe train a multilingual language model with 24 languages with entity representations and show the model consistently outperforms word-based pretrained models in various cross-lingual transfer tasks.
-
15 Jul 2020 4 repositories listedIn this work, we present an information-theoretic framework that formulates cross-lingual language model pre-training as maximizing mutual information between multilingual-multi-granularity texts.
-
24 Mar 2020 4 repositories listed Syntology ran 5 of 10 samples · 5 unverifiedHowever, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse…
-
16 Aug 2019 4 repositories listedRecent years have seen exceptional strides in the task of automatic morphological inflection generation.
-
28 Sep 2021 3 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedThe design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet.
-
20 Jan 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn addition, we demonstrate that we are able to train multilingual word aligners that can obtain robust performance on different language pairs.
-
30 Apr 2020 3 repositories listedThe main goal behind state-of-the-art pre-trained multilingual models such as multilingual BERT and XLM-R is enabling and bootstrapping NLP applications in low-resource languages through zero-shot or few-shot…
-
29 Apr 2020 3 repositories listedWe introduce MultiATIS++, a new multilingual NLU corpus that extends the Multilingual ATIS corpus to nine languages across four language families, and evaluate our method using the corpus.
-
25 Aug 2019 3 repositories listedWe propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.
-
17 Sep 2024 2 repositories listedThis study benchmarks the cross-lingual transfer capabilities from a high-resource language to a low-resource language for both, monolingual and multilingual models, focusing on Kinyarwanda and Kirundi, two Bantu…
-
2 Jul 2024 2 repositories listedWe further present to our knowledge the first use of soft prompts for language transfer, a technique we call soft language prompts.
-
17 Jun 2024 2 repositories listed Syntology ran 9 of 12 samples · 3 unverifiedLarge language models (LLMs) have shown remarkable capabilities in many languages beyond English.
-
14 Jun 2024 2 repositories listedOur approach tackles two essential elements of a language model: the initialization of embeddings and the optimal vocabulary size.
-
26 Mar 2024 2 repositories listedOur findings reveal that (i) current NNRs, even when based on a multilingual language model, suffer from substantial performance losses under ZS-XLT and that (ii) inclusion of target-language data in FS-XLT training has…
-
14 Sep 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedDespite the progress we have recorded in the last few years in multilingual natural language processing, evaluation is typically limited to a small set of languages with available datasets which excludes a large number…
-
31 Aug 2023 2 repositories listed Syntology ran 0 of 6 samples · 6 unverified · 6 pointer-only (licence)We use this dataset to evaluate the capabilities of multilingual masked language models (MLMs) and large language models (LLMs).
-
12 Aug 2023 2 repositories listedCross-lingual open information extraction aims to extract structured information from raw text across multiple languages.
-
3 Apr 2023 2 repositories listedThis paper introduces a Scandinavian benchmarking platform, ScandEval, which can benchmark any pretrained model on four different tasks in the Scandinavian languages.
-
28 Nov 2022 2 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedTranslating training data into many languages has emerged as a practical solution for improving cross-lingual transfer.
-
22 Oct 2022 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedPre-trained multilingual language models show significant performance gains for zero-shot cross-lingual model transfer on a wide range of natural language understanding (NLU) tasks.
-
25 Sep 2022 2 repositories listedWe find that in both settings (legal areas, origin regions), models trained across all groups perform overall better, while they also have improved results in the worst-case scenarios.
-
20 Dec 2021 2 repositories listedLarge-scale generative language models such as GPT-3 are competitive few-shot learners.
-
27 Oct 2021 2 repositories listed Syntology ran 6 of 6 samples · 0 unverifiedWhile recent work on multilingual language models has demonstrated their capacity for cross-lingual zero-shot transfer on downstream tasks, there is a lack of consensus in the community as to what shared properties…
-
14 Oct 2021 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Both these masks can then be composed with the pretrained model.
-
11 Oct 2021 2 repositories listedWav2vec 2.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections