{"url":"/dataset/must-cinema","name":"MuST-Cinema","full_name":null,"description_markdown":"MuST-Cinema is a Multilingual Speech-to-Subtitles corpus ideal for building subtitle-oriented machine and speech translation systems.\r\nIt comprises audio recordings from English TED Talks, which are automatically aligned at the sentence level with their manual transcriptions and translations.\r\n\r\nMuST-Cinema was built by annotating MuST-C with subtitle breaks based on the original subtitle files. Special symbols have been inserted in the aligned sentences to mark subtitle breaks as follows:\r\n\r\n- <eob>: block break (breaks between subtitle blocks)\r\n- <eol>: line breaks (breaks between lines inside the same subtitle block)\r\n\r\nSource: [MuST-Cinema](https://ict.fbk.eu/must-cinema/)","description_withheld":null,"homepage":"https://ict.fbk.eu/must-cinema/","introduced_date":"2020-02-25","introduced_date_note":null,"introduced_by":{"paper":"/paper/must-cinema-a-speech-to-subtitles-corpus","title":"MuST-Cinema: a Speech-to-Subtitles corpus","first_author":"Alina Karakanta","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[],"languages":[],"variants":["MuST-Cinema"],"data_loaders":[],"num_papers_in_archive":12,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}