{"url":"/dataset/speechmatrix","name":"SpeechMatrix","full_name":null,"description_markdown":"**SpeechMatrix** is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of speech.\r\n\r\nSource: [SpeechMatrix: A Large-Scale Mined Corpus\r\nof Multilingual Speech-to-Speech Translations](https://scontent-lhr8-2.xx.fbcdn.net/v/t39.8562-6/310002966_605149234737289_5204270723809834290_n.pdf?_nc_cat=102&ccb=1-7&_nc_sid=ad8a9d&_nc_ohc=FN2KnupyKI0AX90B5UO&_nc_ht=scontent-lhr8-2.xx&oh=00_AT9iFWHchGOnkzVTmwiYIDElIXSnwilSGhDwRQdFh99rlA&oe=63560915)\r\n\r\nImage Source: [(https://scontent-lhr8-2.xx.fbcdn.net/v/t39.8562-6/310002966_605149234737289_5204270723809834290_n.pdf?_nc_cat=102&ccb=1-7&_nc_sid=ad8a9d&_nc_ohc=FN2KnupyKI0AX90B5UO&_nc_ht=scontent-lhr8-2.xx&oh=00_AT9iFWHchGOnkzVTmwiYIDElIXSnwilSGhDwRQdFh99rlA&oe=63560915]((https://scontent-lhr8-2.xx.fbcdn.net/v/t39.8562-6/310002966_605149234737289_5204270723809834290_n.pdf?_nc_cat=102&ccb=1-7&_nc_sid=ad8a9d&_nc_ohc=FN2KnupyKI0AX90B5UO&_nc_ht=scontent-lhr8-2.xx&oh=00_AT9iFWHchGOnkzVTmwiYIDElIXSnwilSGhDwRQdFh99rlA&oe=63560915)","description_withheld":null,"homepage":"https://github.com/facebookresearch/fairseq/tree/ust/examples/speech_matrix","introduced_date":"2022-10-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/speechmatrix-a-large-scale-mined-corpus-of","title":"SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations","first_author":null,"url":null},"license":{"name":"CC-BY-NC 4.0","url":"https://github.com/facebookresearch/fairseq/blob/ust/LICENSE.model"},"modalities":[{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Translation","url":"/task/translation","datasets_with_task":"/datasets/task/translation"},{"name":"Speech-to-Speech Translation","url":"/task/speech-to-speech-translation","datasets_with_task":"/datasets/task/speech-to-speech-translation"}],"languages":[],"variants":["SpeechMatrix"],"data_loaders":[],"num_papers_in_archive":6,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}