{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-distributions-with-linearizing","title":"Predicting distributions with Linearizing Belief Networks","arxiv_id":"1511.05622","date":"2015-11-17","proceeding":null,"authors":["Yann N. Dauphin","David Grangier"],"abstract":"Conditional belief networks introduce stochastic binary variables in neural\nnetworks. Contrary to a classical neural network, a belief network can predict\nmore than the expected value of the output $Y$ given the input $X$. It can\npredict a distribution of outputs $Y$ which is useful when an input can admit\nmultiple outputs whose average is not necessarily a valid answer. Such networks\nare particularly relevant to inverse problems such as image prediction for\ndenoising, or text to speech. However, traditional sigmoid belief networks are\nhard to train and are not suited to continuous problems. This work introduces a\nnew family of networks called linearizing belief nets or LBNs. A LBN decomposes\ninto a deep linear network where each linear unit can be turned on or off by\nnon-deterministic binary latent units. It is a universal approximator of\nreal-valued conditional distributions and can be trained using gradient\ndescent. Moreover, the linear pathways efficiently propagate continuous\ninformation and they act as multiplicative skip-connections that help\noptimization by removing gradient diffusion. This yields a model which trains\nefficiently and improves the state-of-the-art on image denoising and facial\nexpression generation with the Toronto faces dataset.","url_abs":"http://arxiv.org/abs/1511.05622v4","url_pdf":"http://arxiv.org/pdf/1511.05622v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-distributions-with-linearizing","repo_url":"https://github.com/hazdzz/BGU","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"facial-expression-generation","task_name":"Facial expression generation"},{"task_slug":"image-denoising","task_name":"Image Denoising"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"},{"task_slug":null,"task_name":"valid"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}