{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pass-glm-polynomial-approximate-sufficient","title":"PASS-GLM: polynomial approximate sufficient statistics for scalable Bayesian GLM inference","arxiv_id":"1709.09216","date":"2017-09-26","proceeding":"NeurIPS 2017 12","authors":["Jonathan H. Huggins","Ryan P. Adams","Tamara Broderick"],"abstract":"Generalized linear models (GLMs) -- such as logistic regression, Poisson\nregression, and robust regression -- provide interpretable models for diverse\ndata types. Probabilistic approaches, particularly Bayesian ones, allow\ncoherent estimates of uncertainty, incorporation of prior information, and\nsharing of power across experiments via hierarchical models. In practice,\nhowever, the approximate Bayesian methods necessary for inference have either\nfailed to scale to large data sets or failed to provide theoretical guarantees\non the quality of inference. We propose a new approach based on constructing\npolynomial approximate sufficient statistics for GLMs (PASS-GLM). We\ndemonstrate that our method admits a simple algorithm as well as trivial\nstreaming and distributed extensions that do not compound error across\ncomputations. We provide theoretical guarantees on the quality of point (MAP)\nestimates, the approximate posterior, and posterior mean and uncertainty\nestimates. We validate our approach empirically in the case of logistic\nregression using a quadratic approximation and show competitive performance\nwith stochastic gradient descent, MCMC, and the Laplace approximation in terms\nof speed and multiple measures of accuracy -- including on an advertising data\nset with 40 million data points and 20,000 covariates.","url_abs":"http://arxiv.org/abs/1709.09216v3","url_pdf":"http://arxiv.org/pdf/1709.09216v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pass-glm-polynomial-approximate-sufficient","repo_url":"https://bitbucket.org/jhhuggins/pass-glm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"regression-1","task_name":"regression"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.09216","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}