Papers › Efficient-VDVAE: Less is more

Efficient-VDVAE: Less is more

25 Mar 2022arXiv:2203.13751archive 2025-07-28

Louay Hazami, Rayhane Mama, Ragavan Thurairatnam

Hierarchical VAEs have emerged in recent years as a reliable option for maximum likelihood estimation. However, instability issues and demanding computational requirements have hindered research progress in the area. We present simple modifications to the Very Deep VAE to make it converge up to 2.6× faster, save up to 20× in memory load and improve stability during training. Despite these changes, our models achieve comparable or better negative log-likelihood performance than current state-of-the-art models on all $7$ commonly used image datasets we evaluated on. We also make an argument against using 5-bit benchmarks as a way to measure hierarchical VAE's performance due to undesirable biases caused by the 5-bit quantization. Additionally, we empirically demonstrate that roughly 3% of the hierarchical VAE's latent space dimensions is sufficient to encode most of the image information, without loss of performance, opening up the doors to efficiently leverage the hierarchical VAEs' latent space in downstream tasks. We release our source code and models at https://github.com/Rayhane-mamah/Efficient-VDVAE .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Rayhane-mamah/Efficient-VDVAE officialmentioned in papermentioned on GitHubjaxMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationQuantization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Generation Binarized MNIST Efficient-VDVAE nats 79.09 #4 of 10 Archive leaderboard report
Image Generation CelebA 256x256 Efficient-VDVAE bpd 0.51 #1 of 17 Archive leaderboard report
Image Generation CelebA 256x256 Efficient-VDVAE bpd (8-bits) 1.35 #1 of 17 Archive leaderboard report
Image Generation CelebA 64x64 Efficient-VDVAE bits/dimension 1.83 #36 of 39 Archive leaderboard report
Image Generation CelebA-HQ 1024x1024 Efficient-VDVAE bits/dimension 1.01 #10 of 10 Archive leaderboard report
Image Generation FFHQ 1024 x 1024 Efficient-VDVAE bits/dimension 2.30 #19 of 20 Archive leaderboard report
Image Generation FFHQ 256 x 256 Efficient-VDVAE FID 34.88 #40 of 51 Archive leaderboard report
Image Generation FFHQ 256 x 256 Efficient-VDVAE bits/dimension 0.53 #40 of 51 Archive leaderboard report
Image Generation FFHQ 256 x 256 Efficient-VDVAE (DINOv2) FD 514.16 #47 of 51 Archive leaderboard report
Image Generation FFHQ 256 x 256 Efficient-VDVAE (DINOv2) Precision 0.86 #47 of 51 Archive leaderboard report
Image Generation FFHQ 256 x 256 Efficient-VDVAE (DINOv2) Recall 0.14 #47 of 51 Archive leaderboard report
Image Generation ImageNet 64x64 Efficient-VDVAE Bits per dim 3.30 (different downsampling) #37 of 65 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAdaMaxConvolutionCosine AnnealingHierarchical VAEMixture of Logistic DistributionsSigmoid ActivationStochastic Depth

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections