Papers › Enhancing variational generation through self-decomposition

Enhancing variational generation through self-decomposition

6 Feb 2022arXiv:2202.02738archive 2025-07-28

Andrea Asperti, Laura Bugo, Daniele Filippini

In this article we introduce the notion of Split Variational Autoencoder (SVAE), whose output x̂ is obtained as a weighted sum σ⊙x̂₁̂ + (1-σ) ⊙x̂₂̂ of two generated images x̂₁̂,x̂₂̂, and σ is a {\em learned} compositional map. The composing images x̂₁̂,x̂₂̂, as well as the σ-map are automatically synthesized by the model. The network is trained as a usual Variational Autoencoder with a negative loglikelihood loss between training and reconstructed images. No additional loss is required for x̂₁̂,x̂₂̂ or σ, neither any form of human tuning. The decomposition is nondeterministic, but follows two main schemes, that we may roughly categorize as either \say{syntactic} or \say{semantic}. In the first case, the map tends to exploit the strong correlation between adjacent pixels, splitting the image in two complementary high frequency sub-images. In the second case, the map typically focuses on the contours of objects, splitting the image in interesting variations of its content, with more marked and distinctive features. In this case, according to empirical observations, the Fr\'echet Inception Distance (FID) of x̂₁̂ and x̂₂̂ is usually lower (hence better) than that of x̂, that clearly suffers from being the average of the former. In a sense, a SVAE forces the Variational Autoencoder to make choices, in contrast with its intrinsic tendency to {\em average} between alternatives with the aim to minimize the reconstruction loss towards a specific sample. According to the FID metric, our technique, tested on typical datasets such as Mnist, Cifar10 and CelebA, allows us to outperform all previous purely variational architectures (not relying on normalization flows).

PaperPDFCode

Code

asperti/split-vae officialmentioned in papertf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections