Papers › Photorealistic Video Generation with Diffusion Models

Photorealistic Video Generation with Diffusion Models

11 Dec 2023arXiv:2312.06662archive 2025-07-28

Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, José Lezama

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities. Second, for memory and training efficiency, we use a window attention architecture tailored for joint spatial and spatiotemporal generative modeling. Taken together these design decisions enable us to achieve state-of-the-art performance on established video (UCF-101 and Kinetics-600) and image (ImageNet) generation benchmarks without using classifier free guidance. Finally, we also train a cascade of three models for the task of text-to-video generation consisting of a base latent video diffusion model, and two video super-resolution diffusion models to generate videos of 512 ×896 resolution at $8$ frames per second.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Super-ResolutionText-to-Video GenerationVideo GenerationVideo Super-Resolution

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-Video Generation UCF-101 W.A.L.T 3B FVD16 258.1 #3 of 10 Archive leaderboard report
Video Generation Kinetics-600 12 frames, 64x64 W.A.L.T-L FVD 3.3±0.0 #1 of 4 Archive leaderboard report
Video Generation UCF-101 W.A.L.T-XL (class-conditional) FVD16 36±2 #1 of 48 Archive leaderboard report
Video Generation UCF-101 W.A.L.T 3B (text-conditional) FVD16 258.1 #18 of 48 Archive leaderboard report
Video Generation UCF-101 W.A.L.T 3B (text-conditional) Inception Score 35.1 #18 of 48 Archive leaderboard report
Video Prediction Kinetics-600 12 frames, 64x64 W.A.L.T.-L FVD 3.3 #2 of 16 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BASEDiffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections