{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/interpretation-of-machine-learning","title":"Interpretation of machine learning predictions for patient outcomes in electronic health records","arxiv_id":"1903.12074","date":"2019-03-14","proceeding":null,"authors":["William La Cava","Christopher Bauer","Jason H. Moore","Sarah A Pendergrass"],"abstract":"Electronic health records are an increasingly important resource for\nunderstanding the interactions between patient health, environment, and\nclinical decisions. In this paper we report an empirical study of predictive\nmodeling of several patient outcomes using three state-of-the-art machine\nlearning methods. Our primary goal is to validate the models by interpreting\nthe importance of predictors in the final models. Central to interpretation is\nthe use of feature importance scores, which vary depending on the underlying\nmethodology. In order to assess feature importance, we compared univariate\nstatistical tests, information-theoretic measures, permutation testing, and\nnormalized coefficients from multivariate logistic regression models. In\ngeneral we found poor correlation between methods in their assessment of\nfeature importance, even when their performance is comparable and relatively\ngood. However, permutation tests applied to random forest and gradient boosting\nmodels showed the most agreement, and the importance scores matched the\nclinical interpretation most frequently.","url_abs":"http://arxiv.org/abs/1903.12074v1","url_pdf":"http://arxiv.org/pdf/1903.12074v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"interpretation-of-machine-learning","repo_url":"https://github.com/EpistasisLab/interpret_ehr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"feature-importance","task_name":"Feature Importance"}],"methods":[{"method_slug":"logistic-regression","method_name":"Logistic Regression"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}