{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/metboost-exploratory-regression-analysis-with","title":"metboost: Exploratory regression analysis with hierarchically clustered data","arxiv_id":"1702.03994","date":"2017-02-13","proceeding":null,"authors":["Patrick J. Miller","Daniel B. McArtor","Gitta H. Lubke"],"abstract":"As data collections become larger, exploratory regression analysis becomes\nmore important but more challenging. When observations are hierarchically\nclustered the problem is even more challenging because model selection with\nmixed effect models can produce misleading results when nonlinear effects are\nnot included into the model (Bauer and Cai, 2009). A machine learning method\ncalled boosted decision trees (Friedman, 2001) is a good approach for\nexploratory regression analysis in real data sets because it can detect\npredictors with nonlinear and interaction effects while also accounting for\nmissing data. We propose an extension to boosted decision decision trees called\nmetboost for hierarchically clustered data. It works by constraining the\nstructure of each tree to be the same across groups, but allowing the terminal\nnode means to differ. This allows predictors and split points to lead to\ndifferent predictions within each group, and approximates nonlinear group\nspecific effects. Importantly, metboost remains computationally feasible for\nthousands of observations and hundreds of predictors that may contain missing\nvalues. We apply the method to predict math performance for 15,240 students\nfrom 751 schools in data collected in the Educational Longitudinal Study 2002\n(Ingels et al., 2007), allowing 76 predictors to have unique effects for each\nschool. When comparing results to boosted decision trees, metboost has 15%\nimproved prediction performance. Results of a large simulation study show that\nmetboost has up to 70% improved variable selection performance and up to 30%\nimproved prediction performance compared to boosted decision trees when group\nsizes are small","url_abs":"http://arxiv.org/abs/1702.03994v1","url_pdf":"http://arxiv.org/pdf/1702.03994v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"metboost-exploratory-regression-analysis-with","repo_url":"https://github.com/patr1ckm/mvtboost","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"math","task_name":"Math"},{"task_slug":"missing-values","task_name":"Missing Values"},{"task_slug":"model-selection","task_name":"Model Selection"},{"task_slug":"variable-selection","task_name":"Variable Selection"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}