{"url":"/dataset/tr-ar-s2s","name":"TR_AR_S2S","full_name":null,"description_markdown":"Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled equivalents. \r\n\r\nThis work proposes an unsupervised approach to construct speech-to-speech corpus, aligned on short segment levels, to produce a parallel speech corpus in the source- and target- languages. Our methodology exploits video frames, speech recognition, machine translation, and noisy frames removal algorithms to match segments in both languages.","description_withheld":null,"homepage":"https://mailaub-my.sharepoint.com/:u:/g/personal/mab87_mail_aub_edu/EdD4PEjBYb5Khc3hHDhns6kBN32ZvS1OnXztLohWlyT94w?e=LUvdUC","introduced_date":"2022-03-07","introduced_date_note":null,"introduced_by":{"paper":"/paper/creating-speech-to-speech-corpus-from-dubbed","title":"Creating Speech-to-Speech Corpus from Dubbed Series","first_author":"Massa Baali","url":null},"license":null,"modalities":[{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[],"languages":[],"variants":["TR_AR_S2S"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}