{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-conditional-vrnns-for-video","title":"Improved Conditional VRNNs for Video Prediction","arxiv_id":"1904.12165","date":"2019-04-27","proceeding":"ICCV 2019 10","authors":["Lluis Castrejon","Nicolas Ballas","Aaron Courville"],"abstract":"Predicting future frames for a video sequence is a challenging generative\nmodeling task. Promising approaches include probabilistic latent variable\nmodels such as the Variational Auto-Encoder. While VAEs can handle uncertainty\nand model multiple possible future outcomes, they have a tendency to produce\nblurry predictions. In this work we argue that this is a sign of underfitting.\nTo address this issue, we propose to increase the expressiveness of the latent\ndistributions and to use higher capacity likelihood models. Our approach relies\non a hierarchy of latent variables, which defines a family of flexible prior\nand posterior distributions in order to better model the probability of future\nsequences. We validate our proposal through a series of ablation experiments\nand compare our approach to current state-of-the-art latent variable models.\nOur method performs favorably under several metrics in three different\ndatasets.","url_abs":"http://arxiv.org/abs/1904.12165v1","url_pdf":"http://arxiv.org/pdf/1904.12165v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-conditional-vrnns-for-video","repo_url":"https://github.com/facebookresearch/improved_vrnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"video-generation","task_name":"Video Generation"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-generation-on-bair-robot-pushing","task":"Video Generation","dataset":"BAIR Robot Pushing","model":"Hier-VRNN","rank_in_archive_order":16,"of":31,"metrics":{"Cond":"2","FVD score":"143.4","LPIPS":"0.055±0.03","Pred":"28","SSIM":"0.822±0.06","Train":"10"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-bair-robot-pushing","task":"Video Generation","dataset":"BAIR Robot Pushing","model":"VRNN 1L","rank_in_archive_order":18,"of":31,"metrics":{"Cond":"2","FVD score":"149.22","LPIPS":"0.058±0.03","Pred":"28","SSIM":"0.829±0.06","Train":"10"},"uses_additional_data":false},{"leaderboard":"/sota/video-prediction-on-cityscapes-128x128","task":"Video Prediction","dataset":"Cityscapes 128x128","model":"Hier-VRNN","rank_in_archive_order":2,"of":5,"metrics":{"Cond.":"2","FVD":"567.51","LPIPS":"0.264 ± .07","Pred":"28","SSIM":"0.628±0.1","Train":"10"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.12165","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}