{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/coresets-for-scalable-bayesian-logistic","title":"Coresets for Scalable Bayesian Logistic Regression","arxiv_id":"1605.06423","date":"2016-05-20","proceeding":"NeurIPS 2016 12","authors":["Jonathan H. Huggins","Trevor Campbell","Tamara Broderick"],"abstract":"The use of Bayesian methods in large-scale data settings is attractive\nbecause of the rich hierarchical models, uncertainty quantification, and prior\nspecification they provide. Standard Bayesian inference algorithms are\ncomputationally expensive, however, making their direct application to large\ndatasets difficult or infeasible. Recent work on scaling Bayesian inference has\nfocused on modifying the underlying algorithms to, for example, use only a\nrandom data subsample at each iteration. We leverage the insight that data is\noften redundant to instead obtain a weighted subset of the data (called a\ncoreset) that is much smaller than the original dataset. We can then use this\nsmall coreset in any number of existing posterior inference algorithms without\nmodification. In this paper, we develop an efficient coreset construction\nalgorithm for Bayesian logistic regression models. We provide theoretical\nguarantees on the size and approximation quality of the coreset -- both for\nfixed, known datasets, and in expectation for a wide class of data generative\nmodels. Crucially, the proposed approach also permits efficient construction of\nthe coreset in both streaming and parallel settings, with minimal additional\neffort. We demonstrate the efficacy of our approach on a number of synthetic\nand real-world datasets, and find that, in practice, the size of the coreset is\nindependent of the original dataset size. Furthermore, constructing the coreset\ntakes a negligible amount of time compared to that required to run MCMC on it.","url_abs":"http://arxiv.org/abs/1605.06423v3","url_pdf":"http://arxiv.org/pdf/1605.06423v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"coresets-for-scalable-bayesian-logistic","repo_url":"https://bitbucket.org/jhhuggins/lrcoresets","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"coresets-for-scalable-bayesian-logistic","repo_url":"https://github.com/trevorcampbell/bayesian-coresets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"bayesian-inference","task_name":"Bayesian Inference"},{"task_slug":"uncertainty-quantification","task_name":"Uncertainty Quantification"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[{"method_slug":"logistic-regression","method_name":"Logistic Regression"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1605.06423","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}