Home › Datasets › task › Machine Translation
Machine Translation datasets
archive 2025-07-28
82 datasets carry the task tag "Machine Translation" (the task itself: Machine Translation), ordered by the archive's paper count. Page 1 of 2: 48 shown of 82. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Machine Translation datasets 1–48 of 82
WMT 2014 is a collection of datasets used in shared tasks of the Ninth Workshop on Statistical Machine Translation.
288 papers · 9 benchmarks
OpenSubtitles is collection of multilingual parallel corpora.
214 papers · 3 benchmarks
WMT 2016 is a collection of datasets used in shared tasks of the First Conference on Machine Translation.
178 papers · 16 benchmarks
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
Europarl (European Parliament Proceedings Parallel Corpus)
A corpus of parallel text in 21 European languages from the proceedings of the European Parliament.
128 papers · 1 benchmark
FLoRes-200 doubles the existing language coverage of FLoRes-101.
117 papers · 1 benchmark
ASPEC (Asian Scientific Paper Excerpt Corpus)
ASPEC, Asian Scientific Paper Excerpt Corpus, is constructed by the Japan Science and Technology Agency (JST) in collaboration with the National Institute of Information and Communications Technology (NICT).
87 papers · 0 benchmarks
FLoRes-101 is an evaluation benchmark for low-resource and multilingual machine translation.
86 papers · 57 benchmarks
OPUS-100 is an English-centric multilingual corpus covering 100 languages.
78 papers · 0 benchmarks
Europarl-ST is a multilingual Spoken Language Translation corpus containing paired audio-text samples for SLT from and into 9 European languages, for a total of 72 different translation directions.
57 papers · 0 benchmarks
The Shifts Dataset is a dataset for evaluation of uncertainty estimates and robustness to distributional shift.
55 papers · 1 benchmark
The Machine Translation of Noisy Text (MTNT) dataset is a Machine Translation dataset that consists of noisy comments on Reddit and professionally sourced translation.
52 papers · 0 benchmarks
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
Samanantar is the largest publicly available parallel corpora collection for Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu.
40 papers · 0 benchmarks
Tatoeba is a free collection of example sentences with translations geared towards foreign language learners.
38 papers · 2 benchmarks
WMT 2018 is a collection of datasets used in shared tasks of the Third Conference on Machine Translation.
37 papers · 4 benchmarks
WMT 2015 is a collection of datasets used in shared tasks of the Tenth Workshop on Statistical Machine Translation.
33 papers · 2 benchmarks
WMT 2020 is a collection of datasets used in shared tasks of the Fifth Conference on Machine Translation.
33 papers · 0 benchmarks
Consists of millions of entries in which the MT element of the training triplets has been obtained by translating the source side of publicly-available parallel corpora, and using the target side as an artificial human post-edit.
28 papers · 0 benchmarks
News translation is a recurring WMT task.
24 papers · 0 benchmarks
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
MLQE-PE (Multilingual Quality Estimation and Automatic Post-editing Dataset)
The Multilingual Quality Estimation and Automatic Post-editing (MLQE-PE) Dataset is a dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE).
20 papers · 0 benchmarks
PARANMT-50M is a dataset for training paraphrastic sentence embeddings.
13 papers · 0 benchmarks
This is the dataset for the 2020 Duolingo shared task on Simultaneous Translation And Paraphrase for Language Education (STAPLE).
10 papers · 0 benchmarks
A challenging new benchmark for language-agnostic answer retrieval from a multilingual candidate pool.
10 papers · 0 benchmarks
The Japanese-English business conversation corpus, namely Business Scene Dialogue corpus, was constructed in 3 steps: 1.
9 papers · 2 benchmarks
GigaST is a large-scale pseudo speech translation (ST) corpus.
8 papers · 0 benchmarks
The NLC2CMD Competition hosted at NeurIPS 2020 aimed to bring the power of natural language processing to the command line.
8 papers · 1 benchmark
News translation is a recurring WMT task.
8 papers · 0 benchmarks
ACES (A Translation Accuracy Challenge Set)
ACES a dataset consisting of 68 phenomena ranging from simple perturbations at the word/character level to more complex errors based on discourse and real-world knowledge.
7 papers · 1 benchmark
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
Hindi Visual Genome is a multimodal dataset consisting of text and images suitable for English-Hindi multimodal machine translation task and multimodal research.
7 papers · 0 benchmarks
The IWSLT 2017 translation dataset.
7 papers · 1 benchmark
PhoMT is a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs for machine translation.
7 papers · 1 benchmark
BMELD is a bilingual (English-Chinese) dialogue corpus for Neural chat translation.
6 papers · 0 benchmarks
The IWSLT 2015 Evaluation Campaign featured three tracks: automatic speech recognition (ASR), spoken language translation (SLT), and machine translation (MT).
6 papers · 0 benchmarks
JParaCrawl is a parallel corpus for English-Japanese, for which the amount of publicly available parallel corpora is still limited.
6 papers · 0 benchmarks
ShellcodeIA32 is a dataset containing 20 years of shellcodes from a variety of sources is the largest collection of shellcodes in assembly available to date.
6 papers · 1 benchmark
A parallel corpus of Hindi and English, and HindMonoCorp, a monolingual corpus of Hindi in their release version 0.5.
5 papers · 0 benchmarks
MLQE (MultiLingual Quality Estimation)
The MLQE dataset is a dataset for sentence-level Machine Translation Quality Estimation.
5 papers · 0 benchmarks
A new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue.
4 papers · 1 benchmark
FRMT (Few-shot Region-aware Machine Translation)
FRMT is a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation.
4 papers · 4 benchmarks
IndoNLG is a benchmark to measure natural language generation (NLG) progress in three low-resource—yet widely spoken—languages of Indonesia: Indonesian, Javanese, and Sundanese.
4 papers · 0 benchmarks
SRL is the task of extracting semantic predicate-argument structures from sentences.
4 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
HaVG (Hausa Visual Genome Dataset)
A dataset that contains the description of an image or a section within the image in Hausa and its equivalent in English.
3 papers · 0 benchmarks
Itihasa is a large-scale corpus for Sanskrit to English translation containing 93,000 pairs of Sanskrit shlokas and their English translations.
3 papers · 1 benchmark
MMID (Massively Multilingual Image Dataset)
A large-scale multilingual corpus of images, each labeled with the word it represents.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.