Papers › Prompsit's submission to WMT 2018 Parallel Corpus Filtering shared task

Prompsit's submission to WMT 2018 Parallel Corpus Filtering shared task

1 Oct 2018WS 2018 10archive 2025-07-28

V{\'\i}ctor M. S{\'a}nchez-Cartagena, Marta Ba{\~n}{\'o}n, Sergio Ortiz-Rojas, Gema Ram{\'\i}rez

This paper describes Prompsit Language Engineering{'}s submissions to the WMT 2018 parallel corpus filtering shared task. Our four submissions were based on an automatic classifier for identifying pairs of sentences that are mutual translations. A set of hand-crafted hard rules for discarding sentences with evident flaws were applied before the classifier. We explored different strategies for achieving a training corpus with diverse vocabulary and fluent sentences: language model scoring, an active-learning-inspired data selection algorithm and n-gram saturation. Our submissions were very competitive in comparison with other participants on the 100 million word training corpus.

PaperPDFCode

Code

bitextor/bicleaner officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Active LearningLanguage ModelingLanguage ModellingMachine Translation

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections