Datasets › WMT 2019 Metrics Task

WMT 2019 Metrics Task

Introduced by Qingsong Ma et al. in Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges1 Aug 2019 archive 2025-07-28

This shared task will examine automatic evaluation metrics for machine translation. The goals of the shared metrics task are:

To achieve the strongest correlation with human judgement of translation quality; To illustrate the suitability of an automatic evaluation metric as a surrogate for human evaluation; To address problems associated with comparison with a single reference translation; To move automatic evaluation beyond system-level ranking to finer-grained sentence-level ranking.

All datasets for this task are available here.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 3 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • WMT 2019 Metrics Task

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections