Datasets › BUCC
BUCC (Building and Using Comparable Corpora)
The BUCC mining task is a shared task on parallel sentence extraction from two monolingual corpora with a subset of them assumed to be parallel, and that has been available since 2016. For each language pair, the shared task provides a monolingual corpus for each language and a gold mapping list containing true translation pairs. These pairs are the ground truth. The task is to construct a list of translation pairs from the monolingual corpora. The constructed list is compared to the ground truth, and evaluated in terms of the F1 measure.
Source: Language-agnostic BERT Sentence Embedding Image Source: https://comparable.limsi.fr/bucc2017/
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Cross-Lingual Bitext Mining | BUCC French-to-English | Massively Multilingual Sentence Embeddings F1 score 93.91 | Massively Multilingual Sentence Embeddings for Zero-Shot... | facebookresearch/LASER +12 | 3 | Compare |
| Cross-Lingual Bitext Mining | BUCC German-to-English | Massively Multilingual Sentence Embeddings F1 score 96.19 | Massively Multilingual Sentence Embeddings for Zero-Shot... | facebookresearch/LASER +12 | 3 | Compare |
| Cross-Lingual Bitext Mining | BUCC Chinese-to-English | Massively Multilingual Sentence Embeddings F1 score 92.27 | Massively Multilingual Sentence Embeddings for Zero-Shot... | facebookresearch/LASER +12 | 1 | Compare |
| Cross-Lingual Bitext Mining | BUCC Russian-to-English | Massively Multilingual Sentence Embeddings F1 score 93.3 | Massively Multilingual Sentence Embeddings for Zero-Shot... | facebookresearch/LASER +12 | 1 | Compare |
Papers archive 2025-07-28
3 shown of 3 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 42. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond | 13 | 4 | 26 Dec 2018 | ran 4 of 10 samples (6 unverified; 4 pointer-only for licence) |
| Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings | 9 | 2 | 3 Nov 2018 | not harvested |
| Improving Neural Machine Translation Models with Monolingual Data | 2 | 2 | 20 Nov 2015 | ran 6 of 6 samples (0 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- BUCC German-to-English
- BUCC French-to-English
- BUCC Russian-to-English
- BUCC Chinese-to-English
- BUCC
5 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections