Home › Datasets › task › Translation
Translation datasets
archive 2025-07-28
17 datasets carry the task tag "Translation" (the task itself: Translation), ordered by the archive's paper count. Page 1 of 1: 17 shown of 17. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Translation datasets 1–17 of 17
xP3 is a multilingual dataset for multitask prompted finetuning.
34 papers · 0 benchmarks
The Oxford Radar RobotCar Dataset is a radar extension to The Oxford RobotCar Dataset.
28 papers · 4 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
This is the dataset for the 2020 Duolingo shared task on Simultaneous Translation And Paraphrase for Language Education (STAPLE).
10 papers · 0 benchmarks
GigaST is a large-scale pseudo speech translation (ST) corpus.
8 papers · 0 benchmarks
ACES (A Translation Accuracy Challenge Set)
ACES a dataset consisting of 68 phenomena ranging from simple perturbations at the word/character level to more complex errors based on discourse and real-world knowledge.
7 papers · 1 benchmark
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
PhoMT is a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs for machine translation.
7 papers · 1 benchmark
SpeechMatrix is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings.
6 papers · 0 benchmarks
FRMT (Few-shot Region-aware Machine Translation)
FRMT is a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation.
4 papers · 4 benchmarks
HaVG (Hausa Visual Genome Dataset)
A dataset that contains the description of an image or a section within the image in Hausa and its equivalent in English.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
PETCI (PETCI: A Parallel English Translation Dataset of Chinese Idioms)
PETCI is a Parallel English Translation dataset of Chinese Idioms, collected from an idiom dictionary and Google and DeepL translation.
2 papers · 0 benchmarks
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
1 paper · 0 benchmarks
This dataset is parallel text for Bornholmsk and Danish.
1 paper · 0 benchmarks
The FairTranslate Dataset includes 2,418 sentence pairs, each centered around an occupation, designed to assess gender expression and translation in English-French contexts.
1 paper · 0 benchmarks
sPBC (Super Parallel Bible Corpus)
This is a super-parallel Bible corpus containing 1401 language labels (languagescript pairs), meaning that for each verse, we have the translation in other languages.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.