Datasets › WMT 2016

WMT 2016

Introduced by Ond{\v{r}}ej Bojar et al. in Findings of the 2016 Conference on Machine Translation1 Jan 2016 archive 2025-07-28

WMT 2016 is a collection of datasets used in shared tasks of the First Conference on Machine Translation. The conference builds on ten previous Workshops on statistical Machine Translation.

The conference featured ten shared tasks:

  • a news translation task,
  • an IT domain translation task,
  • a biomedical translation task,
  • an automatic post-editing task,
  • a metrics task (assess MT quality given reference translation).
  • a quality estimation task (assess MT quality without access to any reference),
  • a tuning task (optimize a given MT system),
  • a pronoun translation task,
  • a bilingual document alignment task,
  • a multimodal translation task.

Source: http://www.statmt.org/wmt16/index.html

Benchmarks archive 2025-07-28

All 16 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Machine Translation WMT2016 English-Romanian DeLighT BLEU score 34.7 DeLighT: Deep and Light-weight Transformer sacmehta/delight +1 21 Compare
Machine Translation WMT2016 Romanian-English fast-noisy-channel-modeling BLEU score 40.3 Language Models not just for Pre-training: Fast Online... pytorch/fairseq 21 Compare
Machine Translation WMT2016 English-German MADL BLEU score 40.68 Multi-Agent Dual Learning — 12 Compare
Machine Translation WMT2016 German-English FLAN 137B (few-shot, k=11) BLEU score 40.7 Finetuned Language Models Are Zero-Shot Learners hiyouga/llama-efficient-tuning +7 8 Compare
Unsupervised Machine Translation WMT2016 English-German GPT-3 175B (Few-Shot) BLEU 29.7 Language Models are Few-Shot Learners ggml-org/llama.cpp +66 7 Compare
Unsupervised Machine Translation WMT2016 German-English GPT-3 175B (Few-Shot) BLEU 40.6 Language Models are Few-Shot Learners ggml-org/llama.cpp +66 7 Compare
Machine Translation WMT2016 English-Russian Attentional encoder-decoder + BPE BLEU score 26.0 Edinburgh Neural Machine Translation Systems for WMT 16 rsennrich/wmt16-scripts 4 Compare
Unsupervised Machine Translation WMT2016 English-Romanian GPT-3 175B (Few-Shot) BLEU 21 Language Models are Few-Shot Learners ggml-org/llama.cpp +66 3 Compare
Unsupervised Machine Translation WMT2016 Romanian-English GPT-3 175B (Few-Shot) BLEU 39.5 Language Models are Few-Shot Learners ggml-org/llama.cpp +66 3 Compare
Unsupervised Machine Translation WMT2016 English--Romanian BERT-fused NMT BLEU 36.02 Incorporating BERT into Neural Machine Translation bert-nmt/bert-nmt +2 2 Compare
Machine Translation WMT2016 Czech-English Attentional encoder-decoder + BPE BLEU score 31.4 Edinburgh Neural Machine Translation Systems for WMT 16 rsennrich/wmt16-scripts 1 Compare
Machine Translation WMT2016 English-Czech Attentional encoder-decoder + BPE BLEU score 25.8 Edinburgh Neural Machine Translation Systems for WMT 16 rsennrich/wmt16-scripts 1 Compare
Machine Translation WMT2016 English-French DeLighT BLEU score 40.5 DeLighT: Deep and Light-weight Transformer sacmehta/delight +1 1 Compare
Machine Translation WMT2016 Finnish-English CT+B/S construction BLEU 32.4 The University of Sydney's Machine Translation System for WMT19 — 1 Compare
Machine Translation WMT2016 Russian-English Attentional encoder-decoder + BPE BLEU score 28.0 Edinburgh Neural Machine Translation Systems for WMT 16 rsennrich/wmt16-scripts 1 Compare
Binary Classification wmt16 no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 33 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 178. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators 1 1 10 Feb 2024 not harvested
TextBox 2.0: A Text Generation Library with Pre-trained Language Models 1 2 26 Dec 2022 ran 1 of 1 samples (0 unverified)
Finetuned Language Models Are Zero-Shot Learners 8 8 3 Sep 2021 ran 0 of 1 samples (1 unverified)
Language Models not just for Pre-training: Fast Online Neural Noisy Channel Modeling 1 1 13 Nov 2020 not harvested
Incorporating a Local Translation Mechanism into Non-autoregressive Translation 1 4 12 Nov 2020 ran 3 of 5 samples (2 unverified)
Alleviating the Inequality of Attention Heads for Neural Machine Translation 0 2 21 Sep 2020 not harvested
DeLighT: Deep and Light-weight Transformer 2 3 3 Aug 2020 ran 0 of 3 samples (3 unverified)
Language Models are Few-Shot Learners 67 4 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Incorporating BERT into Neural Machine Translation 3 1 17 Feb 2020 not harvested
Exploiting Monolingual Data at Scale for Neural Machine Translation 0 2 1 Nov 2019 not harvested
On the adequacy of untuned warmup for adaptive optimization 1 1 9 Oct 2019 ran 0 of 4 samples (4 unverified)
FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative Flow 2 10 5 Sep 2019 ran 1 of 1 samples (0 unverified)
Adaptively Sparse Transformers 3 2 30 Aug 2019 ran 3 of 3 samples (0 unverified)
The University of Sydney's Machine Translation System for WMT19 0 1 30 Jun 2019 not harvested
Levenshtein Transformer 3 1 27 May 2019 not harvested
MASS: Masked Sequence to Sequence Pre-training for Language Generation 7 4 7 May 2019 not harvested
Multi-Agent Dual Learning 0 1 1 May 2019 not harvested
An Effective Approach to Unsupervised Machine Translation 1 2 4 Feb 2019 not harvested
Cross-lingual Language Model Pretraining 17 6 22 Jan 2019 ran 1 of 7 samples (6 unverified; 1 pointer-only for licence)
Unsupervised Neural Machine Translation with SMT as Posterior Regularization 1 2 14 Jan 2019 ran 0 of 15 samples (15 unverified)
Unsupervised Neural Machine Translation Initialized by Unsupervised Statistical Machine Translation 0 2 30 Oct 2018 not harvested
Unsupervised Statistical Machine Translation 3 2 4 Sep 2018 not harvested
Unsupervised Neural Machine Translation with Weight Sharing 1 2 24 Apr 2018 not harvested
Exploiting Semantics in Neural Machine Translation with Graph Convolutional Networks 0 1 23 Apr 2018 not harvested
Phrase-Based & Neural Unsupervised Machine Translation 14 8 20 Apr 2018 not harvested
Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement 2 2 19 Feb 2018 ran 3 of 4 samples (1 unverified; 1 pointer-only for licence)
Non-Autoregressive Neural Machine Translation 2 2 7 Nov 2017 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Unsupervised Machine Translation Using Monolingual Corpora Only 14 2 31 Oct 2017 ran 3 of 14 samples (11 unverified; 5 pointer-only for licence)
Convolutional Sequence to Sequence Learning 37 1 8 May 2017 not harvested
A Convolutional Encoder Model for Neural Machine Translation 2 2 7 Nov 2016 not harvested

The full list of 33 is in the JSON twin.

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • wmt16 ro-en
  • newstest2016-ende eng-deu
  • newstest2016-deen deu-eng
  • newstest2016-ende
  • newstest2016-deen
  • wmt16
  • WMT2016 En-Ro
  • WMT 2016
  • WMT2016 Russian-English
  • WMT2016 Finnish-English
  • WMT2016 English-Russian
  • WMT2016 English-French
  • WMT2016 English-Czech
  • WMT2016 Czech-English
  • WMT2016 Romanian-English
  • WMT2016 German-English
  • WMT2016 English-Romanian
  • WMT2016 English-German
  • WMT2016 English--Romanian

19 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections