Papers › Photorealistic Video Generation with Diffusion Models
Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, José Lezama
We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities. Second, for memory and training efficiency, we use a window attention architecture tailored for joint spatial and spatiotemporal generative modeling. Taken together these design decisions enable us to achieve state-of-the-art performance on established video (UCF-101 and Kinetics-600) and image (ImageNet) generation benchmarks without using classifier free guidance. Finally, we also train a cascade of three models for the task of text-to-video generation consisting of a base latent video diffusion model, and two video super-resolution diffusion models to generate videos of 512 ×896 resolution at $8$ frames per second.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Text-to-Video Generation | UCF-101 | W.A.L.T 3B | FVD16 | 258.1 | #3 of 10 | Archive leaderboard | report |
| Video Generation | Kinetics-600 12 frames, 64x64 | W.A.L.T-L | FVD | 3.3±0.0 | #1 of 4 | Archive leaderboard | report |
| Video Generation | UCF-101 | W.A.L.T-XL (class-conditional) | FVD16 | 36±2 | #1 of 48 | Archive leaderboard | report |
| Video Generation | UCF-101 | W.A.L.T 3B (text-conditional) | FVD16 | 258.1 | #18 of 48 | Archive leaderboard | report |
| Video Generation | UCF-101 | W.A.L.T 3B (text-conditional) | Inception Score | 35.1 | #18 of 48 | Archive leaderboard | report |
| Video Prediction | Kinetics-600 12 frames, 64x64 | W.A.L.T.-L | FVD | 3.3 | #2 of 16 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections