{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/understanding-random-forests-from-theory-to","title":"Understanding Random Forests: From Theory to Practice","arxiv_id":"1407.7502","date":"2014-07-28","proceeding":null,"authors":["Gilles Louppe"],"abstract":"Data analysis and machine learning have become an integrative part of the\nmodern scientific methodology, offering automated procedures for the prediction\nof a phenomenon based on past observations, unraveling underlying patterns in\ndata and providing insights about the problem. Yet, caution should avoid using\nmachine learning as a black-box tool, but rather consider it as a methodology,\nwith a rational thought process that is entirely dependent on the problem under\nstudy. In particular, the use of algorithms should ideally require a reasonable\nunderstanding of their mechanisms, properties and limitations, in order to\nbetter apprehend and interpret their results.\n  Accordingly, the goal of this thesis is to provide an in-depth analysis of\nrandom forests, consistently calling into question each and every part of the\nalgorithm, in order to shed new light on its learning capabilities, inner\nworkings and interpretability. The first part of this work studies the\ninduction of decision trees and the construction of ensembles of randomized\ntrees, motivating their design and purpose whenever possible. Our contributions\nfollow with an original complexity analysis of random forests, showing their\ngood computational performance and scalability, along with an in-depth\ndiscussion of their implementation details, as contributed within Scikit-Learn.\n  In the second part of this work, we analyse and discuss the interpretability\nof random forests in the eyes of variable importance measures. The core of our\ncontributions rests in the theoretical characterization of the Mean Decrease of\nImpurity variable importance measure, from which we prove and derive some of\nits properties in the case of multiway totally randomized trees and in\nasymptotic conditions. In consequence of this work, our analysis demonstrates\nthat variable importances [...].","url_abs":"http://arxiv.org/abs/1407.7502v3","url_pdf":"http://arxiv.org/pdf/1407.7502v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"understanding-random-forests-from-theory-to","repo_url":"https://github.com/glouppe/phd-thesis","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"understanding-random-forests-from-theory-to","repo_url":"https://github.com/ysraell/random-forest-lab","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1407.7502","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}