Papers › Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource Settings

Model and Data Transfer for Cross-Lingual Sequence Labelling in Zero-Resource Settings

23 Oct 2022arXiv:2210.12623archive 2025-07-28

Iker García-Ferrero, Rodrigo Agerri, German Rigau

Zero-resource cross-lingual transfer approaches aim to apply supervised models from a source language to unlabelled target languages. In this paper we perform an in-depth study of the two main techniques employed so far for cross-lingual zero-resource sequence labelling, based either on data or model transfer. Although previous research has proposed translation and annotation projection (data-based cross-lingual transfer) as an effective technique for cross-lingual sequence labelling, in this paper we experimentally demonstrate that high capacity multilingual language models applied in a zero-shot (model-based cross-lingual transfer) setting consistently outperform data-based cross-lingual transfer approaches. A detailed analysis of our results suggests that this might be due to important differences in language use. More specifically, machine translation often generates a textual signal which is different to what the models are exposed to when using gold standard data, which affects both the fine-tuning and evaluation processes. Our results also indicate that data-based cross-lingual transfer approaches remain a competitive option when high-capacity multilingual language models are not available.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ikergarcia1996/annotation-projection-app officialmentioned in papermentioned on GitHub report
ikergarcia1996/easy-label-projection officialmentioned in papermentioned on GitHubpytorch report
ikergarcia1996/Easy-Translate officialmentioned in paperpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Lingual NERCross-Lingual TransferMachine TranslationTranslation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Lingual NER CoNLL 2003 XLM-RoBERTa-large Dutch 82.3 #1 of 4 Archive leaderboard report
Cross-Lingual NER CoNLL 2003 XLM-RoBERTa-large German 74.5 #1 of 4 Archive leaderboard report
Cross-Lingual NER CoNLL 2003 XLM-RoBERTa-large Spanish 79.5 #1 of 4 Archive leaderboard report
Cross-Lingual NER CoNLL Dutch XLM-R large F1 79.7 #7 of 10 Archive leaderboard report
Cross-Lingual NER CoNLL German XLM-R large F1 74.5 #4 of 10 Archive leaderboard report
Cross-Lingual NER CoNLL Spanish XLM-R large F1 79.5 #1 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections