{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/preventing-posterior-collapse-with-delta-vaes","title":"Preventing Posterior Collapse with delta-VAEs","arxiv_id":"1901.03416","date":"2019-01-10","proceeding":"ICLR 2019 5","authors":["Ali Razavi","Aäron van den Oord","Ben Poole","Oriol Vinyals"],"abstract":"Due to the phenomenon of \"posterior collapse,\" current latent variable\ngenerative models pose a challenging design choice that either weakens the\ncapacity of the decoder or requires augmenting the objective so it does not\nonly maximize the likelihood of the data. In this paper, we propose an\nalternative that utilizes the most powerful generative models as decoders,\nwhilst optimising the variational lower bound all while ensuring that the\nlatent variables preserve and encode useful information. Our proposed\n$\\delta$-VAEs achieve this by constraining the variational family for the\nposterior to have a minimum distance to the prior. For sequential latent\nvariable models, our approach resembles the classic representation learning\napproach of slow feature analysis. We demonstrate the efficacy of our approach\nat modeling text on LM1B and modeling images: learning representations,\nimproving sample quality, and achieving state of the art log-likelihood on\nCIFAR-10 and ImageNet $32\\times 32$.","url_abs":"http://arxiv.org/abs/1901.03416v1","url_pdf":"http://arxiv.org/pdf/1901.03416v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-generation-on-imagenet-32x32","task":"Image Generation","dataset":"ImageNet 32x32","model":"δ-VAE","rank_in_archive_order":22,"of":35,"metrics":{"bpd":"3.77"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1901.03416","atlas_url":"https://app.syntology.ai/?focus=1901.03416","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}