Papers › JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

1 Jul 2019ACL 2019 7archive 2025-07-28

{\v{Z}}eljko Agi{\'c}, Ivan Vuli{\'c}

Viable cross-lingual transfer critically depends on the availability of parallel texts. Shortage of such resources imposes a development and evaluation bottleneck in multilingual processing. We introduce JW300, a parallel corpus of over 300 languages with around 100 thousand parallel sentences per language pair on average. In this paper, we present the resource and showcase its utility in experiments with cross-lingual word embedding induction and multi-source part-of-speech projection.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Lingual Transfer

Datasets

Introduced by this paper, per the archive.

JW300

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections