Papers › Stochastic Variational Video Prediction
Stochastic Variational Video Prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, Sergey Levine
Predicting the future in real-world settings, particularly from raw sensory observations such as images, is exceptionally challenging. Real-world events can be stochastic and unpredictable, and the high dimensionality and complexity of natural images requires the predictive model to build an intricate understanding of the natural world. Many existing methods tackle this problem by making simplifying assumptions about the environment. One common assumption is that the outcome is deterministic and there is only one plausible future. This can lead to low-quality predictions in real-world settings with stochastic dynamics. In this paper, we develop a stochastic variational video prediction (SV2P) method that predicts a different possible future for each sample of its latent variables. To the best of our knowledge, our model is the first to provide effective stochastic multi-frame prediction for real-world video. We demonstrate the capability of the proposed method in predicting detailed future frames of videos on multiple real-world datasets, both action-free and action-conditioned. We find that our proposed method produces substantially improved video predictions when compared to the same model without stochasticity, and to other stochastic video prediction methods. Our SV2P implementation will be open sourced upon publication.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video Generation | BAIR Robot Pushing | SV2P (from FVD) | Cond | 2 | #25 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from FVD) | FVD score | 262.5 | #25 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from FVD) | Pred | 14 | #25 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from FVD) | Train | 14 | #25 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | Cond | 2 | #30 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | FVD score | 965±17 | #30 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | LPIPS | 0.0912±0.0053 | #30 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | PSNR | 20.39±0.27 | #30 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | Pred | 28 | #30 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | SSIM | 0.8169±0.0086 | #30 of 31 | Archive leaderboard | report |
| Video Generation | BAIR Robot Pushing | SV2P (from SRVP) | Train | 12 | #30 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Cond | 10 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | FVD | 209.5 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | LPIPS | 0.232 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | PSNR | 25.87 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Params (M) | 8.3 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Pred | 40 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | SSIM | 0.782 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Train | 10 | #5 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Cond | 10 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | FVD | 253.5 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | LPIPS | 0.260 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | PSNR | 25.70 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Params (M) | 8.3 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Pred | 40 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | SSIM | 0.772 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P time-invariant (from Grid-keypoints) | Train | 10 | #8 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | Cond | 10 | #12 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | FVD | 636 ± 1 | #12 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | LPIPS | 0.2049±0.0053 | #12 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | PSNR | 28.19±0.31 | #12 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | Pred | 30 | #12 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | SSIM | 0.838 | #12 of 31 | Archive leaderboard | report |
| Video Prediction | KTH | SV2P (from SRVP) | Train | 10 | #12 of 31 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections