Datasets › BUCC

BUCC (Building and Using Comparable Corpora)

Introduced by Pierre Zweigenbaum et al. in Overview of the Second BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora1 Jan 2017 archive 2025-07-28

The BUCC mining task is a shared task on parallel sentence extraction from two monolingual corpora with a subset of them assumed to be parallel, and that has been available since 2016. For each language pair, the shared task provides a monolingual corpus for each language and a gold mapping list containing true translation pairs. These pairs are the ground truth. The task is to construct a list of translation pairs from the monolingual corpora. The constructed list is compared to the ground truth, and evaluated in terms of the F1 measure.

Source: Language-agnostic BERT Sentence Embedding Image Source: https://comparable.limsi.fr/bucc2017/

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Cross-Lingual Bitext Mining BUCC French-to-English Massively Multilingual Sentence Embeddings F1 score 93.91 Massively Multilingual Sentence Embeddings for Zero-Shot... facebookresearch/LASER +12 3 Compare
Cross-Lingual Bitext Mining BUCC German-to-English Massively Multilingual Sentence Embeddings F1 score 96.19 Massively Multilingual Sentence Embeddings for Zero-Shot... facebookresearch/LASER +12 3 Compare
Cross-Lingual Bitext Mining BUCC Chinese-to-English Massively Multilingual Sentence Embeddings F1 score 92.27 Massively Multilingual Sentence Embeddings for Zero-Shot... facebookresearch/LASER +12 1 Compare
Cross-Lingual Bitext Mining BUCC Russian-to-English Massively Multilingual Sentence Embeddings F1 score 93.3 Massively Multilingual Sentence Embeddings for Zero-Shot... facebookresearch/LASER +12 1 Compare

Papers archive 2025-07-28

3 shown of 3 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 42. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond 13 4 26 Dec 2018 ran 4 of 10 samples (6 unverified; 4 pointer-only for licence)
Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings 9 2 3 Nov 2018 not harvested
Improving Neural Machine Translation Models with Monolingual Data 2 2 20 Nov 2015 ran 6 of 6 samples (0 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • BUCC German-to-English
  • BUCC French-to-English
  • BUCC Russian-to-English
  • BUCC Chinese-to-English
  • BUCC

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections