{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stochastic-variational-video-prediction","title":"Stochastic Variational Video Prediction","arxiv_id":"1710.11252","date":"2017-10-30","proceeding":"ICLR 2018 1","authors":["Mohammad Babaeizadeh","Chelsea Finn","Dumitru Erhan","Roy H. Campbell","Sergey Levine"],"abstract":"Predicting the future in real-world settings, particularly from raw sensory\nobservations such as images, is exceptionally challenging. Real-world events\ncan be stochastic and unpredictable, and the high dimensionality and complexity\nof natural images requires the predictive model to build an intricate\nunderstanding of the natural world. Many existing methods tackle this problem\nby making simplifying assumptions about the environment. One common assumption\nis that the outcome is deterministic and there is only one plausible future.\nThis can lead to low-quality predictions in real-world settings with stochastic\ndynamics. In this paper, we develop a stochastic variational video prediction\n(SV2P) method that predicts a different possible future for each sample of its\nlatent variables. To the best of our knowledge, our model is the first to\nprovide effective stochastic multi-frame prediction for real-world video. We\ndemonstrate the capability of the proposed method in predicting detailed future\nframes of videos on multiple real-world datasets, both action-free and\naction-conditioned. We find that our proposed method produces substantially\nimproved video predictions when compared to the same model without\nstochasticity, and to other stochastic video prediction methods. Our SV2P\nimplementation will be open sourced upon publication.","url_abs":"http://arxiv.org/abs/1710.11252v2","url_pdf":"http://arxiv.org/pdf/1710.11252v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stochastic-variational-video-prediction","repo_url":"https://github.com/RoboTurk-Platform/roboturk_real_dataset","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"stochastic-variational-video-prediction","repo_url":"https://github.com/StanfordVL/roboturk_real_dataset","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"stochastic-variational-video-prediction","repo_url":"https://github.com/suraj-nair-1/google-research/tree/master/hierarchical_foresight","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"video-generation","task_name":"Video Generation"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-generation-on-bair-robot-pushing","task":"Video Generation","dataset":"BAIR Robot Pushing","model":"SV2P (from FVD)","rank_in_archive_order":25,"of":31,"metrics":{"Cond":"2","FVD score":"262.5","Pred":"14","Train":"14"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-bair-robot-pushing","task":"Video Generation","dataset":"BAIR Robot Pushing","model":"SV2P (from SRVP)","rank_in_archive_order":30,"of":31,"metrics":{"Cond":"2","FVD score":"965±17","LPIPS":"0.0912±0.0053","PSNR":"20.39±0.27","Pred":"28","SSIM":"0.8169±0.0086","Train":"12"},"uses_additional_data":false},{"leaderboard":"/sota/video-prediction-on-kth","task":"Video Prediction","dataset":"KTH","model":"SV2P time-invariant (from Grid-keypoints)","rank_in_archive_order":5,"of":31,"metrics":{"Cond":"10","FVD":"209.5","LPIPS":"0.232","PSNR":"25.87","Params (M)":"8.3","Pred":"40","SSIM":"0.782","Train":"10"},"uses_additional_data":false},{"leaderboard":"/sota/video-prediction-on-kth","task":"Video Prediction","dataset":"KTH","model":"SV2P time-invariant (from Grid-keypoints)","rank_in_archive_order":8,"of":31,"metrics":{"Cond":"10","FVD":"253.5","LPIPS":"0.260","PSNR":"25.70","Params (M)":"8.3","Pred":"40","SSIM":"0.772","Train":"10"},"uses_additional_data":false},{"leaderboard":"/sota/video-prediction-on-kth","task":"Video Prediction","dataset":"KTH","model":"SV2P (from SRVP)","rank_in_archive_order":12,"of":31,"metrics":{"Cond":"10","FVD":"636 ± 1","LPIPS":"0.2049±0.0053","PSNR":"28.19±0.31","Pred":"30","SSIM":"0.838","Train":"10"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.11252","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}