{"url":"/dataset/wmt19-metrics-task","name":"WMT 2019 Metrics Task","full_name":null,"description_markdown":"This shared task will examine automatic evaluation metrics for machine translation. The goals of the shared metrics task are:\r\n\r\nTo achieve the strongest correlation with human judgement of translation quality;\r\nTo illustrate the suitability of an automatic evaluation metric as a surrogate for human evaluation;\r\nTo address problems associated with comparison with a single reference translation;\r\nTo move automatic evaluation beyond system-level ranking to finer-grained sentence-level ranking.\r\n\r\nAll datasets for this task are available [here](http://www.statmt.org/wmt19/metrics-task.html).","description_withheld":null,"homepage":"http://www.statmt.org/wmt19/metrics-task.html","introduced_date":"2019-08-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/results-of-the-wmt19-metrics-shared-task","title":"Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges","first_author":"Qingsong Ma","url":null},"license":{"name":"Unknown","url":null},"modalities":[],"tasks":[],"languages":[],"variants":["WMT 2019 Metrics Task"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}