{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/parallel-gaussian-process-regression-for-big","title":"Parallel Gaussian Process Regression for Big Data: Low-Rank Representation Meets Markov Approximation","arxiv_id":"1411.4510","date":"2014-11-17","proceeding":null,"authors":["Kian Hsiang Low","Jiangbo Yu","Jie Chen","Patrick Jaillet"],"abstract":"The expressive power of a Gaussian process (GP) model comes at a cost of poor\nscalability in the data size. To improve its scalability, this paper presents a\nlow-rank-cum-Markov approximation (LMA) of the GP model that is novel in\nleveraging the dual computational advantages stemming from complementing a\nlow-rank approximate representation of the full-rank GP based on a support set\nof inputs with a Markov approximation of the resulting residual process; the\nlatter approximation is guaranteed to be closest in the Kullback-Leibler\ndistance criterion subject to some constraint and is considerably more refined\nthan that of existing sparse GP models utilizing low-rank representations due\nto its more relaxed conditional independence assumption (especially with larger\ndata). As a result, our LMA method can trade off between the size of the\nsupport set and the order of the Markov property to (a) incur lower\ncomputational cost than such sparse GP models while achieving predictive\nperformance comparable to them and (b) accurately represent features/patterns\nof any scale. Interestingly, varying the Markov order produces a spectrum of\nLMAs with PIC approximation and full-rank GP at the two extremes. An advantage\nof our LMA method is that it is amenable to parallelization on multiple\nmachines/cores, thereby gaining greater scalability. Empirical evaluation on\nthree real-world datasets in clusters of up to 32 computing nodes shows that\nour centralized and parallel LMA methods are significantly more time-efficient\nand scalable than state-of-the-art sparse and full-rank GP regression methods\nwhile achieving comparable predictive performances.","url_abs":"http://arxiv.org/abs/1411.4510v1","url_pdf":"http://arxiv.org/pdf/1411.4510v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"parallel-gaussian-process-regression-for-big","repo_url":"https://github.com/arikcj/pgpr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"regression-1","task_name":"regression"}],"methods":[{"method_slug":"gaussian-process","method_name":"Gaussian Process"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1411.4510","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}