{"url":"/dataset/wmt-2015-news","name":"WMT 2015 News","full_name":"WMT 2015 News Translation Task","description_markdown":"News translation is a recurring WMT task. The test set is a collection of parallel corpora consisting of about 1500 English sentences translated into 5 languages (Czech, German, Finnish, French, Russian) and additional 1500 sentences from each of the 5 languages translated to English. The sentences are taken from newspaper articles for each language pair, except for French, where the test set was drawn from user-generated comments on the news articles (from Guardian and Le Monde). The translation was done by professional translators.\n\nThe training data consists of parallel corpora to train translation models, monolingual corpora to train language models and development sets for tuning.\nSome training corpora were identical from WMT 2014 (Europarl, United Nations, French-English 10⁹ corpus, CzEng, Common Crawl, Russian-English parallel data provided by Yandex, Wikipedia Headlines provided by CMU) and some were update (News Commentary, monolingual news data). Additionally, the Finnish Europarl and Finnish-English Wikipedia Headline corpus were added.\n\nSource: [https://paperswithcode.com/paper/findings-of-the-2016-conference-on-machine/](https://paperswithcode.com/paper/findings-of-the-2016-conference-on-machine/)\nImage Source: [httpshttps://www.aclweb.org/anthology/W15-3001.pdf](httpshttps://www.aclweb.org/anthology/W15-3001.pdf)","description_withheld":null,"homepage":"http://www.statmt.org/wmt15/index.html","introduced_date":"2015-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/findings-of-the-2015-workshop-on-statistical","title":"Findings of the 2015 Workshop on Statistical Machine Translation","first_author":"Ond{\\v{r}}ej Bojar","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Parallel","url":"/datasets/modality/parallel"}],"tasks":[{"name":"Machine Translation","url":"/task/machine-translation","datasets_with_task":"/datasets/task/machine-translation"},{"name":"Automatic Post-Editing","url":"/task/automatic-post-editing","datasets_with_task":"/datasets/task/automatic-post-editing"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"},{"name":"German","url":"/datasets/language/german"},{"name":"Russian","url":"/datasets/language/russian"},{"name":"Czech","url":"/datasets/language/czech"},{"name":"Finnish","url":"/datasets/language/finnish"}],"variants":["WMT 2015 News"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}