{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/finding-structure-in-data-using-multivariate","title":"Finding structure in data using multivariate tree boosting","arxiv_id":"1511.02025","date":"2015-11-06","proceeding":null,"authors":["Patrick J. Miller","Gitta H. Lubke","Daniel B. McArtor","C. S. Bergeman"],"abstract":"Technology and collaboration enable dramatic increases in the size of\npsychological and psychiatric data collections, but finding structure in these\nlarge data sets with many collected variables is challenging. Decision tree\nensembles like random forests (Strobl, Malley, and Tutz, 2009) are a useful\ntool for finding structure, but are difficult to interpret with multiple\noutcome variables which are often of interest in psychology. To find and\ninterpret structure in data sets with multiple outcomes and many predictors\n(possibly exceeding the sample size), we introduce a multivariate extension to\na decision tree ensemble method called Gradient Boosted Regression Trees\n(Friedman, 2001). Our method, multivariate tree boosting, can be used for\nidentifying important predictors, detecting predictors with non-linear effects\nand interactions without specification of such effects, and for identifying\npredictors that cause two or more outcome variables to covary without\nparametric assumptions. We provide the R package 'mvtboost' to estimate, tune,\nand interpret the resulting model, which extends the implementation of\nunivariate boosting in the R package 'gbm' (Ridgeway, 2013) to continuous,\nmultivariate outcomes. To illustrate the approach, we analyze predictors of\npsychological well-being (Ryff and Keyes, 1995). Simulations verify that our\napproach identifies predictors with non-linear effects and achieves high\nprediction accuracy, exceeding or matching the performance of (penalized)\nmultivariate multiple regression and multivariate decision trees over a wide\nrange of conditions.","url_abs":"http://arxiv.org/abs/1511.02025v2","url_pdf":"http://arxiv.org/pdf/1511.02025v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"finding-structure-in-data-using-multivariate","repo_url":"https://github.com/patr1ckm/mvtboost","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}