Home › Datasets › task › Machine Translation
Machine Translation datasets
archive 2025-07-28
82 datasets carry the task tag "Machine Translation" (the task itself: Machine Translation), ordered by the archive's paper count. Page 2 of 2: 34 shown of 82. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Machine Translation datasets 49–82 of 82
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
APE (Automatic Post-Editing)
APE is useful to evaluate Machine Translation automatic post-editing (APE), which is the task of improving the output of a blackbox MT system by automatically fixing its mistakes.
2 papers · 0 benchmarks
Amharic - English Parallel Corpus for Machine Translation contains 33,955 sentence pairs extracted text from such news platforms as Ethiopian Press Agency1, Fana Broadcasting Corporate2, and Walta Information Center3.
2 papers · 0 benchmarks
A multilingual etymological database extracted from the Wiktionary (described in Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0)
2 papers · 0 benchmarks
Given a sentence in the source language, generate a translation in the target language.
2 papers · 0 benchmarks
GATITOS (Google's Additional Translations Into Tail-languages: Often Short)
The GATITOS (Google's Additional Translations Into Tail-languages: Often Short) dataset is a high-quality, multi-way parallel dataset of tokens and short phrases, intended for training and improving machine translation models.
2 papers · 0 benchmarks
The MUSE dataset contains bilingual dictionaries for 110 pairs of languages.
2 papers · 2 benchmarks
PETCI (PETCI: A Parallel English Translation Dataset of Chinese Idioms)
PETCI is a Parallel English Translation dataset of Chinese Idioms, collected from an idiom dictionary and Google and DeepL translation.
2 papers · 0 benchmarks
Perseus is a dataset for Cross-Lingual Summarization (CLS) which collects about 94K Chinese scientific documents paired with English summaries.
2 papers · 0 benchmarks
SubEdits is a human-annnoated post-editing dataset of neural machine translation outputs, compiled from in-house NMT outputs and human post-edits of subtitles form Rakuten Viki.
2 papers · 0 benchmarks
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
The Alexa Point of View dataset is point of view conversion dataset, a parallel corpus of messages spoken to a virtual assistant and the converted messages for delivery.
1 paper · 1 benchmark
BPersona-chat is an evaluation dataset based on the English multiturn chat corpus Persona-chat and the Japanese multiturn chat corpus JPersona-chat.
1 paper · 0 benchmarks
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
1 paper · 0 benchmarks
A bilingual corpus of English-Arabic parallel tweets and a list of Twitter accounts who post English-Arabic tweets regularly.
1 paper · 0 benchmarks
French sentences are sourced from Tatoeba repository and then translated into Congolese Swahili.
1 paper · 0 benchmarks
FGraDA (Fine-Grained Domain Adaptation Dataset)
Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real- world…
1 paper · 0 benchmarks
HumanMT is a collection of human ratings and corrections of machine translations.
1 paper · 0 benchmarks
IgboNLP is a standard machine translation benchmark dataset for Igbo.
1 paper · 0 benchmarks
a parallel corpus of Sorani (ckb or Central Kurdish) and Kurmanji (kmr or Northern Kurdish) dialects of Kurdish along with English (eng).
1 paper · 0 benchmarks
MENYO-20k is the first multi-domain parallel corpus with a special focus on clean orthography for Yorùbá--English with standardized train-test splits for benchmarking.
1 paper · 0 benchmarks
Dataset Description The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag.
1 paper · 1 benchmark
PheMT is a phenomenon-wise dataset designed for evaluating the robustness of Japanese-English machine translation systems.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
SimpEvalASSET is a dataset for learning learnable metrics using modern language models.
1 paper · 0 benchmarks
The Vashantor dataset consists of 32,500 sentences from different regions, including Chittagong, Noakhali, Sylhet, Barishal, and Mymensingh.
1 paper · 0 benchmarks
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
News translation is a recurring WMT task.
1 paper · 0 benchmarks
The Biomedical Translation Shared Task was first introduced at the First Conference of Machine Translation.
1 paper · 0 benchmarks
The IT Translation Task is a shared task introduced in the First Conference on Machine Translation.
1 paper · 0 benchmarks
scb-mt-en-th-2020 is an English-Thai machine translation dataset with over 1 million segment pairs, curated from various sources, namely news, Wikipedia articles, SMS messages, task-based dialogs, web-crawled data and government documents.
1 paper · 0 benchmarks
ContraCAT (Contrastive Coreference Analytical Templates (for Machine Translation))
Current approaches to context-aware MT rely on a set of surface heuristics to translate pronouns, which break down when translations require real reasoning.
0 papers · 0 benchmarks
WMT 2021 Ge'ez-Amharic is a Ge'ez-Amharic dataset prepared for NMT tasks of the 6th Workshop on NLP at Debre Berhan University, Ethiopia.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.