Datasets › MuST-C
MuST-C
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation. It covers eight language directions, from English to German, Spanish, French, Italian, Dutch, Portuguese, Romanian and Russian. The corpus consists of audio, transcriptions and translations of English TED talks, and it comes with a predefined training, validation and test split.
Source: One-to-Many Multilingual End-to-End Speech Translation Image Source: https://mt.fbk.eu/must-c
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Speech-to-Text Translation | MuST-C EN->DE | Task Modulation + Multitask Learning(ASR/MT) + Data Augmentation Case-sensitive sacreBLEU 28.88 | TASK AWARE MULTI-TASK LEARNING FOR SPEECH TO TEXT TASKS | — | 8 | Compare |
| Speech-to-Text Translation | MuST-C | Transformer with Adapters SacreBLEU 26.61 | Lightweight Adapter Tuning for Multilingual Speech Translation | formiel/fairseq +1 | 2 | Compare |
Papers archive 2025-07-28
8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 216. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Speechformer: Reducing Information Loss in Direct Speech Translation | 1 | 1 | 9 Sep 2021 | ran 2 of 5 samples (3 unverified; 4 pointer-only for licence) |
| TASK AWARE MULTI-TASK LEARNING FOR SPEECH TO TEXT TASKS | 0 | 1 | 10 Jun 2021 | not harvested |
| Lightweight Adapter Tuning for Multilingual Speech Translation | 2 | 2 | 2 Jun 2021 | not harvested |
| End-to-End Speech Translation with Pre-trained Models and Adapters: UPC at IWSLT 2021 | 1 | 1 | 10 May 2021 | not harvested |
| NeurST: Neural Speech Translation Toolkit | 1 | 1 | 18 Dec 2020 | not harvested |
| Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation | 1 | 2 | 2 Nov 2020 | not harvested |
| fairseq S2T: Fast Speech-to-Text Modeling with fairseq | 5 | 1 | 11 Oct 2020 | not harvested |
| End-to-End Offline Speech Translation System for IWSLT 2020 using Modality Agnostic Meta-Learning | 0 | 1 | 1 Jul 2020 | not harvested |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- MuST-C EN->DE
- MuST-C
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections