{"url":"/dataset/storytelling-video-dataset","name":"Video Dataset","full_name":"Storytelling Video Dataset (Russian, Emotion, Gesture, Speech)","description_markdown":"The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers. Each video is 10+ minutes long and includes synchronized speech, facial expressions, gestures, and emotional variation. The dataset is ideal for research and development in:\r\n\r\n🗣️ Speech-to-text & voice modeling (Russian)\r\n\r\n😃 Emotion & gesture recognition\r\n\r\n🤖 Multimodal learning & LLM alignment\r\n\r\n🧍 Avatar generation & digital human training\r\n\r\nKey Features:\r\n\r\nFull-body framing (waist-up or knee-up)\r\n\r\nClear facial expressions and gestures\r\n\r\nHigh-quality speech audio in Russian\r\n\r\nNatural storytelling with emotional diversity\r\n\r\nPreview (10 Participants):\r\nWe’ve prepared a short compilation of 10 different speakers from the dataset. Each clip is taken from a unique 10-minute unscripted video, featuring full-body framing, gestures, and emotional speech.\r\n\r\n📺 Watch the sample:\r\nhttps://drive.google.com/drive/folders/1QtC2il-Qb62nZNJlOtG8WvOf-SSkSGM-?usp=drive_link – Video Preview\r\n\r\nDataset Pages:\r\n\r\n🐙 GitHub: github.com/MaratDV\r\n\r\n📊 DataHub: datahub.io/@MaratDV/storytelling-video-dataset\r\n\r\nemail chinzad@gmail.com or telegram @Marat_DV","description_withheld":null,"homepage":"https://huggingface.co/datasets/MaratDV/video-dataset","introduced_date":"2025-05-01","introduced_date_note":null,"introduced_by":null,"license":{"name":"other (custom license)","url":"https://huggingface.co/datasets/MaratDV/video-dataset#license"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Emotion Recognition","url":"/task/emotion-recognition","datasets_with_task":"/datasets/task/emotion-recognition"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"},{"name":"Video Classification","url":"/task/video-classification","datasets_with_task":"/datasets/task/video-classification"},{"name":"Multimodal Emotion Recognition","url":"/task/multimodal-emotion-recognition","datasets_with_task":"/datasets/task/multimodal-emotion-recognition"}],"languages":[{"name":"Russian","url":"/datasets/language/russian"}],"variants":["Video Dataset"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}