Datasets › Video Dataset

Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))

1 May 2025 archive 2025-07-28

The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers. Each video is 10+ minutes long and includes synchronized speech, facial expressions, gestures, and emotional variation. The dataset is ideal for research and development in:

🗣️ Speech-to-text & voice modeling (Russian)

😃 Emotion & gesture recognition

🤖 Multimodal learning & LLM alignment

🧍 Avatar generation & digital human training

Key Features:

Full-body framing (waist-up or knee-up)

Clear facial expressions and gestures

High-quality speech audio in Russian

Natural storytelling with emotional diversity

Preview (10 Participants): We’ve prepared a short compilation of 10 different speakers from the dataset. Each clip is taken from a unique 10-minute unscripted video, featuring full-body framing, gestures, and emotional speech.

📺 Watch the sample: https://drive.google.com/drive/folders/1QtC2il-Qb62nZNJlOtG8WvOf-SSkSGM-?usp=drive_link – Video Preview

Dataset Pages:

🐙 GitHub: github.com/MaratDV

📊 DataHub: datahub.io/@MaratDV/storytelling-video-dataset

email chinzad@gmail.com or telegram @Marat_DV

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

other (custom license)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Video Dataset

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections