Home › Datasets › task › Video Generation
Video Generation datasets
archive 2025-07-28
24 datasets carry the task tag "Video Generation" (the task itself: Video Generation), ordered by the archive's paper count. Page 1 of 1: 24 shown of 24. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Video Generation datasets 1–24 of 24
UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
WebVid contains 10 million video clips with captions, sourced from the web.
257 papers · 1 benchmark
LAION-400M is a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.
169 papers · 1 benchmark
The Kinetics-600 is a large-scale action recognition dataset which consists of around 480K videos from 600 action categories.
148 papers · 3 benchmarks
Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes.
95 papers · 3 benchmarks
InternVid is a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodAL understanding and generation.
46 papers · 0 benchmarks
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
Dataset of 64x64 images of a robot pushing objects on a table top.
27 papers · 2 benchmarks
CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy.
15 papers · 0 benchmarks
OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data.
8 papers · 1 benchmark
YouTube Driving Dataset contains a massive amount of real-world driving frames with various conditions, from different weather, different regions, to diverse scene types
6 papers · 1 benchmark
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
The Deep Fakes Dataset is a collection of "in the wild" portrait videos for deepfake detection.
3 papers · 0 benchmarks
QST contains 1,167 video clips that are cut out from 216 time-lapse 4K videos collected from YouTube, which can be used for a variety of tasks, such as (high-resolution) video generation, (high-resolution) video prediction,…
3 papers · 0 benchmarks
AVSync15 is a high-quality synchronized audio-video dataset curated from VGGSound.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
DropletVideo is a project exploring high-order spatio-temporal consistency in image-to-video generation.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
Description - Repository: Code, Page, Data - Paper: arxiv.org/abs/2411.17440 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a star and citation.
1 paper · 0 benchmarks
🏃♂️ Open-HypermotionX Dataset Open-Hypermotion is a large-scale, high-quality dataset designed for training and evaluating pose-guided human image animation models, with a special focus on complex, dynamic human motions (Hypermotion),…
1 paper · 0 benchmarks
We create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples.
1 paper · 1 benchmark
TLFM dataset (TLFM dataset for microscopy image sequence generation)
TLFM dataset structured in sequences of at least nine timesteps.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.