Home › Datasets › task › Cross-Lingual Transfer
Cross-Lingual Transfer datasets
archive 2025-07-28
15 datasets carry the task tag "Cross-Lingual Transfer" (the task itself: Cross-Lingual Transfer), ordered by the archive's paper count. Page 1 of 1: 15 shown of 15. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Cross-Lingual Transfer datasets 1–15 of 15
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
The Cross-lingual Choice of Plausible Alternatives (XCOPA) dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across languages.
110 papers · 1 benchmark
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
Tatoeba is a free collection of example sentences with translations geared towards foreign language learners.
38 papers · 2 benchmarks
TyDiQA is the gold passage version of the Typologically Diverse Question Answering (TyDiWA) dataset, a benchmark for information-seeking question answering, which covers nine languages.
38 papers · 1 benchmark
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
JESC (Japanese-English Subtitle Corpus)
Japanese-English Subtitle Corpus is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue.
17 papers · 0 benchmarks
A large multilingual benchmark, XL-WiC, featuring gold standards in 12 new languages from varied language families and with different degrees of resource availability, opening room for evaluation scenarios such as zero-shot cross-lingual…
15 papers · 0 benchmarks
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.