{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/orthogonally-decoupled-variational-gaussian","title":"Orthogonally Decoupled Variational Gaussian Processes","arxiv_id":"1809.08820","date":"2018-09-24","proceeding":"NeurIPS 2018 12","authors":["Hugh Salimbeni","Ching-An Cheng","Byron Boots","Marc Deisenroth"],"abstract":"Gaussian processes (GPs) provide a powerful non-parametric framework for\nreasoning over functions. Despite appealing theory, its superlinear\ncomputational and memory complexities have presented a long-standing challenge.\nState-of-the-art sparse variational inference methods trade modeling accuracy\nagainst complexity. However, the complexities of these methods still scale\nsuperlinearly in the number of basis functions, implying that that sparse GP\nmethods are able to learn from large datasets only when a small model is used.\nRecently, a decoupled approach was proposed that removes the unnecessary\ncoupling between the complexities of modeling the mean and the covariance\nfunctions of a GP. It achieves a linear complexity in the number of mean\nparameters, so an expressive posterior mean function can be modeled. While\npromising, this approach suffers from optimization difficulties due to\nill-conditioning and non-convexity. In this work, we propose an alternative\ndecoupled parametrization. It adopts an orthogonal basis in the mean function\nto model the residues that cannot be learned by the standard coupled approach.\nTherefore, our method extends, rather than replaces, the coupled approach to\nachieve strictly better performance. This construction admits a straightforward\nnatural gradient update rule, so the structure of the information manifold that\nis lost during decoupling can be leveraged to speed up learning. Empirically,\nour algorithm demonstrates significantly faster convergence in multiple\nexperiments.","url_abs":"http://arxiv.org/abs/1809.08820v3","url_pdf":"http://arxiv.org/pdf/1809.08820v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"orthogonally-decoupled-variational-gaussian","repo_url":"https://github.com/hughsalimbeni/orth_decoupled_var_gps","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"gaussian-processes","task_name":"Gaussian Processes"},{"task_slug":"variational-inference","task_name":"Variational Inference"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1809.08820","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}