{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/when-can-multi-site-datasets-be-pooled-for","title":"When can Multi-Site Datasets be Pooled for Regression? Hypothesis Tests, $\\ell_2$-consistency and Neuroscience Applications","arxiv_id":"1709.00640","date":"2017-09-02","proceeding":"ICML 2017 8","authors":["Hao Henry Zhou","Yilin Zhang","Vamsi K. Ithapu","Sterling C. Johnson","Grace Wahba","Vikas Singh"],"abstract":"Many studies in biomedical and health sciences involve small sample sizes due\nto logistic or financial constraints. Often, identifying weak (but\nscientifically interesting) associations between a set of predictors and a\nresponse necessitates pooling datasets from multiple diverse labs or groups.\nWhile there is a rich literature in statistical machine learning to address\ndistributional shifts and inference in multi-site datasets, it is less clear\n${\\it when}$ such pooling is guaranteed to help (and when it does not) --\nindependent of the inference algorithms we use. In this paper, we present a\nhypothesis test to answer this question, both for classical and high\ndimensional linear regression. We precisely identify regimes where pooling\ndatasets across multiple sites is sensible, and how such policy decisions can\nbe made via simple checks executable on each site before any data transfer ever\nhappens. With a focus on Alzheimer's disease studies, we present empirical\nresults showing that in regimes suggested by our analysis, pooling a local\ndataset with data from an international study improves power.","url_abs":"http://arxiv.org/abs/1709.00640v1","url_pdf":"http://arxiv.org/pdf/1709.00640v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"when-can-multi-site-datasets-be-pooled-for","repo_url":"https://github.com/hzhoustat/ICML2017","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.00640","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}