{"url":"/dataset/dailymoth-70h","name":"DailyMoth-70h","full_name":null,"description_markdown":"DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news channel [TheDailyMoth](https://www.youtube.com/c/TheDailyMoth). The primary purpose of the dataset is to be used as a benchmark and analysis dataset for (gloss-free) sign language translation.\r\n\r\nThe dataset comes with four parts:\r\n\r\n- **raw_videos**: Contains the unsegmented DailyMoth videos (496 in total) with blurring applied to the burnt-in captions and advertisement breaks and news headline banners\r\n\r\n- **blurred_clips**: Contains the segmented video clips (48386 in total) with facial blurring applied. Each clip comes in its native frame rate (either 24, 29.97 or 30 fps) and as 224x224px region-of-interest (ROI) crops around the signer\r\n\r\n- **unblurred_clips**: Contains the unblurred segmented video clips (48386 in total). Each clip comes in its native frame rate (either 24, 29.97 or 30 fps) and as 224x224px region-of-interest (ROI) crops around the signer\r\n\r\n- **manifests**: Contains the manifest TSV files for training, validation, and testing. Also contains a combined manifest file and a tsv file with the start and end timestamps used to segment the raw video\r\n\r\nDetailed dataset statistics are listed in [https://arxiv.org/abs/2402.09611](https://arxiv.org/abs/2402.09611).\r\n\r\nThe dataset is available for download at [https://github.com/facebookresearch/ssvp_slt?tab=readme-ov-file#dailymoth-70h](https://github.com/facebookresearch/ssvp_slt?tab=readme-ov-file#dailymoth-70h).","description_withheld":null,"homepage":"https://github.com/facebookresearch/ssvp_slt","introduced_date":"2024-02-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/towards-privacy-aware-sign-language","title":"Towards Privacy-Aware Sign Language Translation at Scale","first_author":"Phillip Rust","url":null},"license":{"name":"CC BY-NC 4.0","url":"https://github.com/facebookresearch/ssvp_slt/blob/main/LICENSE"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sign Language Translation","url":"/task/sign-language-translation","datasets_with_task":"/datasets/task/sign-language-translation"},{"name":"Gloss-free Sign Language Translation","url":"/task/gloss-free-sign-language-translation","datasets_with_task":"/datasets/task/gloss-free-sign-language-translation"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"American Sign Language","url":"/datasets/language/american-sign-language"}],"variants":["DailyMoth-70h"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}