Papers › Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings

Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings

3 Nov 2018ACL 2019 7arXiv:1811.01136archive 2025-07-28

Mikel Artetxe, Holger Schwenk

Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for this task based on multilingual sentence embeddings. In contrast to previous approaches, which rely on nearest neighbor retrieval with a hard threshold over cosine similarity, our proposed method accounts for the scale inconsistencies of this measure, considering the margin between a given sentence pair and its closest candidates instead. Our experiments show large improvements over existing methods. We outperform the best published results on the BUCC mining task and the UN reconstruction task by more than 10 F1 and 30 precision points, respectively. Filtering the English-German ParaCrawl corpus with our approach, we obtain 31.2 BLEU points on newstest2014, an improvement of more than one point over the best official filtered version.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

facebookresearch/LASER officialmentioned in papermentioned on GitHubpytorch report
LawrenceDuan/myLASER mentioned on GitHubpytorch report
Tony4469/laser-agir mentioned on GitHubpytorchNOASSERTION report
imamathcat/LASER_Dependencies mentioned on GitHubpytorchNOASSERTION report
kmkwon94/ainize-laser mentioned on GitHubpytorchNOASSERTION report
prabhakar267/LASER-improved mentioned on GitHubpytorchNOASSERTION report
raymondhs/fairseq-laser mentioned on GitHubpytorchMIT report
thompsonb/prism_bitext_filter mentioned on GitHubpytorch report
transducens/LASERtrain mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Lingual Bitext MiningMachine TranslationParallel Corpus MiningRetrievalSentenceSentence EmbeddingsTranslation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Lingual Bitext Mining BUCC French-to-English Multilingual Sentence Embeddings F1 score 92.89 #2 of 3 Archive leaderboard report
Cross-Lingual Bitext Mining BUCC German-to-English Multilingual Sentence Embeddings F1 score 95.58 #2 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections