Browse State-of-the-Art › Video Prediction
Video Prediction
208 papers with code · 19 benchmarks · 25 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
19 leaderboard tables shown for this task, 19 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 19 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
25 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 208 papers with code (394 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Jun 2015 23 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 2 pointer-only (licence)The goal of precipitation nowcasting is to predict the future rainfall intensity in a local region over a relatively short period of time.
-
25 May 2016 17 repositories listed Syntology ran 0 of 6 samples · 6 unverified · 2 pointer-only (licence)Here, we explore prediction of future frames in a video sequence as an unsupervised learning rule for learning about the structure of the visual world.
-
17 Apr 2018 11 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedWe present PredRNN++, an improved recurrent network for video predictive learning.
-
20 Aug 2018 10 repositories listedWe study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.
-
7 Nov 2022 7 repositories listed Syntology ran 12 of 15 samples · 3 unverifiedNotably, MogaNet hits 80.
-
7 Apr 2022 5 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedGenerating temporally coherent high fidelity video is an important milestone in generative modeling research.
-
13 Jun 2017 5 repositories listedNeural networks trained on datasets such as ImageNet have led to major advances in visual object classification.
-
17 Nov 2015 5 repositories listed Syntology ran 2 of 10 samples · 8 unverifiedLearning to predict future images from a video sequence involves the construction of an internal representation that models the image evolution accurately, and therefore, to some degree, its content and dynamics.
-
5 Jan 2020 4 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 2 pointer-only (licence)The adjoint sensitivity method scalably computes gradients of solutions to ordinary differential equations.
-
1 Oct 2019 4 repositories listedInspired by the success of sparse motion-based prediction for video compression, we propose a parametric video prediction on a sparse motion field composed of few critical pixels and their motion vectors.
-
19 Nov 2018 4 repositories listedNatural spatiotemporal processes can be highly non-stationary in many ways, e.
-
4 Apr 2018 4 repositories listed Syntology ran 4 of 16 samples · 12 unverified · 3 pointer-only (licence)However, learning to predict raw future observations, such as frames in a video, is exceedingly challenging -- the ambiguous nature of the problem can cause a naively designed model to average together possible futures…
-
12 Jun 2017 4 repositories listed Syntology ran 9 of 10 samples · 1 unverified · 5 pointer-only (licence)To address these problems, we propose both a new model and a benchmark for precipitation nowcasting.
-
3 Aug 2016 4 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedComma.
-
14 Mar 2024 3 repositories listed Syntology ran 6 of 8 samples · 2 unverifiedIn this paper, we introduce the first large-scale video prediction model in the autonomous driving discipline.
-
14 Feb 2024 3 repositories listedExponential Moving Average (EMA) is a widely used weight averaging (WA) regularization to learn flat optima for better generalizations without extra cost in deep neural network (DNN) optimization.
-
9 Oct 2023 3 repositories listed Syntology ran 12 of 20 samples · 8 unverifiedWhile Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation.
-
23 May 2023 3 repositories listed Syntology ran 6 of 15 samples · 9 unverified · 1 pointer-only (licence)A promising approach is to extract preferences for behaviors from unlabeled videos, which are widely available on the internet.
-
9 Jun 2022 3 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 9 pointer-only (licence)From CNN, RNN, to ViT, we have witnessed remarkable advancements in video prediction, incorporating auxiliary inputs, elaborate neural architectures, and sophisticated training strategies.
-
17 Mar 2021 3 repositories listedThis paper models these structures by presenting PredRNN, a new recurrent network, in which a pair of memory cells are explicitly decoupled, operate in nearly independent transition manners, and finally form unified…
-
3 Mar 2020 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedLeveraging physical knowledge described by partial differential equations (PDEs) is an appealing way to improve unsupervised video prediction methods.
-
14 Jun 2019 3 repositories listedWe fill in the gap by evaluating PredNet both as an implementation of the predictive coding theory and as a self-supervised video prediction model using a challenging video action classification dataset.
-
1 May 2019 3 repositories listedWe first evaluate the E3D-LSTM network on widely-used future video prediction datasets and achieve the state-of-the-art performance.
-
3 Dec 2018 3 repositories listedTo this extent we propose Fr\'{e}chet Video Distance (FVD), a new metric for generative models of video, and StarCraft 2 Videos (SCV), a benchmark of game play from custom starcraft 2 scenarios that challenge the…
-
6 Oct 2018 3 repositories listedWe demonstrate that this idea can be combined with a video-prediction based controller to enable complex behaviors to be learned from scratch using only raw visual inputs, including grasping, repositioning objects, and…
-
21 Feb 2018 3 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Sample generations are both varied and sharp, even many frames into the future, and compare favorably to those from existing approaches.
-
30 Oct 2017 3 repositories listedWe find that our proposed method produces substantially improved video predictions when compared to the same model without stochasticity, and to other stochastic video prediction methods.
-
15 Oct 2017 3 repositories listedOne learning signal that is always available for autonomously collected data is prediction: if a robot can learn to predict the future, it can use this predictive model to take actions to produce desired outcomes, such…
-
8 Feb 2017 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We combine the advantages of these two methods by training a deep network that learns to synthesize video frames by flowing pixel values from existing ones, which we call deep voxel flow.
-
4 Dec 2023 2 repositories listedTo address this challenge, we present the Super-Multivariate Urban Mobility Transformer (SUMformer), which utilizes a specially designed attention mechanism to calculate temporal and cross-variable correlations and…
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections