Home › Datasets › task › Video Prediction

Video Prediction datasets

archive 2025-07-28

25 datasets carry the task tag "Video Prediction" (the task itself: Video Prediction), ordered by the archive's paper count. Page 1 of 1: 25 shown of 25. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Video Prediction datasets 1–25 of 25

The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The Human3.6M dataset is one of the largest motion capture datasets, which consists of 3.6 million human poses and corresponding images captured by a high-speed motion capture system.
783 papers · 13 benchmarks
DAVIS (Densely Annotated VIdeo Segmentation)
The Densely Annotation Video Segmentation dataset (DAVIS) is a high quality and high resolution densely annotated video segmentation dataset under two resolutions, 480p and 1080p.
734 papers · 10 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
KTH (KTH Action dataset)
The efforts to create a non-trivial and publicly available dataset for action recognition was initiated at the KTH Royal Institute of Technology in 2004.
279 papers · 2 benchmarks
The Vimeo-90K is a large-scale high-quality video dataset for lower-level video processing.
220 papers · 3 benchmarks
MPI (Max Planck Institute) Sintel is a dataset for optical flow evaluation that has 1064 synthesized stereo images and ground truth data for disparity.
198 papers · 5 benchmarks
The Moving MNIST dataset contains 10,000 video sequences, each consisting of 20 frames.
194 papers · 1 benchmark
The Kinetics-600 is a large-scale action recognition dataset which consists of around 480K videos from 600 action categories.
148 papers · 3 benchmarks
The YouTube-8M dataset is a large scale video dataset, which includes more than 7 million videos with 4716 classes labeled by the annotation system.
147 papers · 2 benchmarks
Sprites (2D Video Game Character Sprites)
The Sprites dataset contains 60 pixel color images of animated characters (sprites).
52 papers · 3 benchmarks
PHYRE (PHYsical REasoning)
Benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical environment.
35 papers · 2 benchmarks
Dataset of 64x64 images of a robot pushing objects on a table top.
27 papers · 2 benchmarks
The Robotic Pushing Dataset is a dataset for video prediction for real-world interactive agents which consists of 59,000 robot interactions involving pushing motions, including a test set with novel objects.
12 papers · 0 benchmarks
EarthNet2021 (EarthNet2021: Earth Surface Forecasting)
Satellite images are snapshots of the Earth surface.
10 papers · 4 benchmarks
SynPick is a synthetic dataset for dynamic scene understanding in bin-picking scenarios.
7 papers · 1 benchmark
QST (Quick Sky Time)
QST contains 1,167 video clips that are cut out from 216 time-lapse 4K videos collected from YouTube, which can be used for a variety of tasks, such as (high-resolution) video generation, (high-resolution) video prediction,…
3 papers · 0 benchmarks
CloudCast (CloudCast: A Satellite-Based Dataset and Baseline for Forecasting Clouds)
A satellite-based dataset called "CloudCast".
1 paper · 0 benchmarks
IISc VINE (Indian Institute of Science VIdeo Naturalness Evaluation)
Indian Institute of Science VIdeo Naturalness Evaluation (IISc VINE) is a database consisting of 300 videos, obtained by applying different prediction models on different datasets, and accompanying human opinion scores.
1 paper · 0 benchmarks
A parameterized synthetic dataset called Moving Symbols to support the objective study of video prediction networks.
1 paper · 0 benchmarks
Shanghai2020 (Shanghai-2020 Dataset)
It is released by the Shanghai Central Meteorological Observatory (SCMO) in 2020, records serval years of historical precipitation events in the Yangtze River delta area.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.