{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bayesian-dark-knowledge","title":"Bayesian Dark Knowledge","arxiv_id":"1506.04416","date":"2015-06-14","proceeding":"NeurIPS 2015 12","authors":["Anoop Korattikara","Vivek Rathod","Kevin Murphy","Max Welling"],"abstract":"We consider the problem of Bayesian parameter estimation for deep neural\nnetworks, which is important in problem settings where we may have little data,\nand/ or where we need accurate posterior predictive densities, e.g., for\napplications involving bandits or active learning. One simple approach to this\nis to use online Monte Carlo methods, such as SGLD (stochastic gradient\nLangevin dynamics). Unfortunately, such a method needs to store many copies of\nthe parameters (which wastes memory), and needs to make predictions using many\nversions of the model (which wastes time).\n  We describe a method for \"distilling\" a Monte Carlo approximation to the\nposterior predictive density into a more compact form, namely a single deep\nneural network. We compare to two very recent approaches to Bayesian neural\nnetworks, namely an approach based on expectation propagation [Hernandez-Lobato\nand Adams, 2015] and an approach based on variational Bayes [Blundell et al.,\n2015]. Our method performs better than both of these, is much simpler to\nimplement, and uses less computation at test time.","url_abs":"http://arxiv.org/abs/1506.04416v3","url_pdf":"http://arxiv.org/pdf/1506.04416v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bayesian-dark-knowledge","repo_url":"https://github.com/ofnt/real_anot","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"parameter-estimation","task_name":"parameter estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1506.04416","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}