Papers › Efficient-VDVAE: Less is more
Efficient-VDVAE: Less is more
Louay Hazami, Rayhane Mama, Ragavan Thurairatnam
Hierarchical VAEs have emerged in recent years as a reliable option for maximum likelihood estimation. However, instability issues and demanding computational requirements have hindered research progress in the area. We present simple modifications to the Very Deep VAE to make it converge up to 2.6× faster, save up to 20× in memory load and improve stability during training. Despite these changes, our models achieve comparable or better negative log-likelihood performance than current state-of-the-art models on all $7$ commonly used image datasets we evaluated on. We also make an argument against using 5-bit benchmarks as a way to measure hierarchical VAE's performance due to undesirable biases caused by the 5-bit quantization. Additionally, we empirically demonstrate that roughly 3% of the hierarchical VAE's latent space dimensions is sufficient to encode most of the image information, without loss of performance, opening up the doors to efficiently leverage the hierarchical VAEs' latent space in downstream tasks. We release our source code and models at https://github.com/Rayhane-mamah/Efficient-VDVAE .
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Image Generation | Binarized MNIST | Efficient-VDVAE | nats | 79.09 | #4 of 10 | Archive leaderboard | report |
| Image Generation | CelebA 256x256 | Efficient-VDVAE | bpd | 0.51 | #1 of 17 | Archive leaderboard | report |
| Image Generation | CelebA 256x256 | Efficient-VDVAE | bpd (8-bits) | 1.35 | #1 of 17 | Archive leaderboard | report |
| Image Generation | CelebA 64x64 | Efficient-VDVAE | bits/dimension | 1.83 | #36 of 39 | Archive leaderboard | report |
| Image Generation | CelebA-HQ 1024x1024 | Efficient-VDVAE | bits/dimension | 1.01 | #10 of 10 | Archive leaderboard | report |
| Image Generation | FFHQ 1024 x 1024 | Efficient-VDVAE | bits/dimension | 2.30 | #19 of 20 | Archive leaderboard | report |
| Image Generation | FFHQ 256 x 256 | Efficient-VDVAE | FID | 34.88 | #40 of 51 | Archive leaderboard | report |
| Image Generation | FFHQ 256 x 256 | Efficient-VDVAE | bits/dimension | 0.53 | #40 of 51 | Archive leaderboard | report |
| Image Generation | FFHQ 256 x 256 | Efficient-VDVAE (DINOv2) | FD | 514.16 | #47 of 51 | Archive leaderboard | report |
| Image Generation | FFHQ 256 x 256 | Efficient-VDVAE (DINOv2) | Precision | 0.86 | #47 of 51 | Archive leaderboard | report |
| Image Generation | FFHQ 256 x 256 | Efficient-VDVAE (DINOv2) | Recall | 0.14 | #47 of 51 | Archive leaderboard | report |
| Image Generation | ImageNet 64x64 | Efficient-VDVAE | Bits per dim | 3.30 (different downsampling) | #37 of 65 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections