{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fast-yet-simple-natural-gradient-descent-for","title":"Fast yet Simple Natural-Gradient Descent for Variational Inference in Complex Models","arxiv_id":"1807.04489","date":"2018-07-12","proceeding":null,"authors":["Mohammad Emtiyaz Khan","Didrik Nielsen"],"abstract":"Bayesian inference plays an important role in advancing machine learning, but\nfaces computational challenges when applied to complex models such as deep\nneural networks. Variational inference circumvents these challenges by\nformulating Bayesian inference as an optimization problem and solving it using\ngradient-based optimization. In this paper, we argue in favor of\nnatural-gradient approaches which, unlike their gradient-based counterparts,\ncan improve convergence by exploiting the information geometry of the\nsolutions. We show how to derive fast yet simple natural-gradient updates by\nusing a duality associated with exponential-family distributions. An attractive\nfeature of these methods is that, by using natural-gradients, they are able to\nextract accurate local approximations for individual model components. We\nsummarize recent results for Bayesian deep learning showing the superiority of\nnatural-gradient approaches over their gradient counterparts.","url_abs":"http://arxiv.org/abs/1807.04489v2","url_pdf":"http://arxiv.org/pdf/1807.04489v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fast-yet-simple-natural-gradient-descent-for","repo_url":"https://github.com/ssggreg/active_learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"bayesian-inference","task_name":"Bayesian Inference"},{"task_slug":"variational-inference","task_name":"Variational Inference"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.04489","atlas_url":"https://app.syntology.ai/?focus=1807.04489","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}