{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/imagen-video-high-definition-video-generation","title":"Imagen Video: High Definition Video Generation with Diffusion Models","arxiv_id":"2210.02303","date":"2022-10-05","proceeding":null,"authors":["Jonathan Ho","William Chan","Chitwan Saharia","Jay Whang","Ruiqi Gao","Alexey Gritsenko","Diederik P. Kingma","Ben Poole","Mohammad Norouzi","David J. Fleet","Tim Salimans"],"abstract":"We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models. Given a text prompt, Imagen Video generates high definition videos using a base video generation model and a sequence of interleaved spatial and temporal video super-resolution models. We describe how we scale up the system as a high definition text-to-video model including design decisions such as the choice of fully-convolutional temporal and spatial super-resolution models at certain resolutions, and the choice of the v-parameterization of diffusion models. In addition, we confirm and transfer findings from previous work on diffusion-based image generation to the video generation setting. Finally, we apply progressive distillation to our video models with classifier-free guidance for fast, high quality sampling. We find Imagen Video not only capable of generating videos of high fidelity, but also having a high degree of controllability and world knowledge, including the ability to generate diverse videos and text animations in various artistic styles and with 3D object understanding. See https://imagen.research.google/video/ for samples.","url_abs":"https://arxiv.org/abs/2210.02303v1","url_pdf":"https://arxiv.org/pdf/2210.02303v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"super-resolution","task_name":"Super-Resolution"},{"task_slug":"video-generation","task_name":"Video Generation"},{"task_slug":"video-super-resolution","task_name":"Video Super-Resolution"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":"world-knowledge","task_name":"World Knowledge"}],"methods":[{"method_slug":"base","method_name":"BASE"},{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-generation-on-laion-400m","task":"Video Generation","dataset":"LAION-400M","model":"Imagen original (constant=6)","rank_in_archive_order":1,"of":6,"metrics":{"CLIP":"25.19","CLIP R-Precision":"92.12"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-laion-400m","task":"Video Generation","dataset":"LAION-400M","model":"Imagen fully distilled (oscillate (15,1))","rank_in_archive_order":2,"of":6,"metrics":{"CLIP R-Precision":"90.97"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-laion-400m","task":"Video Generation","dataset":"LAION-400M","model":"Imagen distilled (constant=6)","rank_in_archive_order":3,"of":6,"metrics":{"CLIP":"25.29","CLIP R-Precision":"90.88"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-laion-400m","task":"Video Generation","dataset":"LAION-400M","model":"Imagen original (oscillate(15,1))","rank_in_archive_order":4,"of":6,"metrics":{"CLIP":"25.03","CLIP R-Precision":"89.91"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-laion-400m","task":"Video Generation","dataset":"LAION-400M","model":"Imagen fully distilled (constant=6)","rank_in_archive_order":5,"of":6,"metrics":{"CLIP R-Precision":"89.68"},"uses_additional_data":false},{"leaderboard":"/sota/video-generation-on-laion-400m","task":"Video Generation","dataset":"LAION-400M","model":"Imagen distilled (oscillate (15,1))","rank_in_archive_order":6,"of":6,"metrics":{"CLIP":"25.12","CLIP R-Precision":"88.78"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2210.02303","atlas_url":"https://app.syntology.ai/?focus=2210.02303","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}