Browse State-of-the-Art › Video Prediction

Video Prediction

208 papers with code · 19 benchmarks · 25 datasets archive 2025-07-28

Computer VisionTime Series

Benchmarks archive 2025-07-28

19 leaderboard tables shown for this task, 19 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 19 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
KTH (31 rows) Grid-keypoints Accurate Grid Keypoint Learning for Efficient Video Prediction code Syntology ran 2 of 2 samples · 0 unverified Compare
Moving MNIST (31 rows) PredFormer Video Prediction Transformers without Recurrence or Convolution code — Compare
Kinetics-600 12 frames, 64x64 (16 rows) SiD2 Simpler Diffusion (SiD2): 1.5 FID on ImageNet512 with pixel-space diffusion — — Compare
Human3.6M (9 rows) IAM4VP Implicit Stacked Autoregressive Model for Video Prediction code — Compare
BAIR Robot Pushing (6 rows) MAGVIT (-L-FP) MAGVIT: Masked Generative Video Transformer code Syntology ran 1 of 10 samples · 9 unverified Compare
Cityscapes 128x128 (5 rows) GHVAEs Greedy Hierarchical Variational Autoencoders for Large-Scale Video... — — Compare
SynpickVP (5 rows) MSPred MSPred: Video Prediction at Multiple Spatio-Temporal Scales with... code — Compare
CMU Mocap-2 (4 rows) Latent SDE Scalable Gradients for Stochastic Differential Equations code Syntology ran 1 of 5 samples · 4 unverified Compare
Cityscapes (3 rows) DMVFN A Dynamic Multi-Scale Voxel Flow Network for Video Prediction code Syntology ran 3 of 4 samples · 1 unverified Compare
KITTI (3 rows) DMVFN A Dynamic Multi-Scale Voxel Flow Network for Video Prediction code Syntology ran 3 of 4 samples · 1 unverified Compare
Vimeo90K (3 rows) OPT — — — Compare
CMU Mocap-1 (2 rows) ODE2VAE-KL ODE²VAE: Deep generative second order ODEs with Bayesian neural networks code Syntology ran 0 of 4 samples · 4 unverified Compare
DAVIS 2017 (2 rows) DMVFN A Dynamic Multi-Scale Voxel Flow Network for Video Prediction code Syntology ran 3 of 4 samples · 1 unverified Compare
Colored dSprites (1 row) MGP-VAE (with geodesic loss) Disentangling Multiple Features in Video Sequences using Gaussian... code — Compare
KTH 64x64 cond10 pred30 (1 row) SRVP Stochastic Latent Residual Video Prediction code Syntology ran 2 of 12 samples · 10 unverified Compare
MPI Sintel (1 row) MCnet [villegas2017mcnet] Temporal View Synthesis of Dynamic Scenes through 3D Object Motion... code — Compare
Something-Something V2 (1 row) MAGVIT MAGVIT: Masked Generative Video Transformer code Syntology ran 1 of 10 samples · 9 unverified Compare
Sprites (1 row) MGP-VAE (with geodesic loss) Disentangling Multiple Features in Video Sequences using Gaussian... code — Compare
YouTube-8M (1 row) SDCNet SDC-Net: Video prediction using spatially-displaced convolution code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

25 datasets whose archive record lists this task, ordered by the archive's paper count.

Subtasks archive 2025-07-28

2 subtasks in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 208 papers with code (394 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections