{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalising-random-forest-parameter","title":"Generalising Random Forest Parameter Optimisation to Include Stability and Cost","arxiv_id":"1706.09865","date":"2017-06-29","proceeding":null,"authors":["C. H. Bryan Liu","Benjamin Paul Chamberlain","Duncan A. Little","Angelo Cardoso"],"abstract":"Random forests are among the most popular classification and regression\nmethods used in industrial applications. To be effective, the parameters of\nrandom forests must be carefully tuned. This is usually done by choosing values\nthat minimize the prediction error on a held out dataset. We argue that error\nreduction is only one of several metrics that must be considered when\noptimizing random forest parameters for commercial applications. We propose a\nnovel metric that captures the stability of random forests predictions, which\nwe argue is key for scenarios that require successive predictions. We motivate\nthe need for multi-criteria optimization by showing that in practical\napplications, simply choosing the parameters that lead to the lowest error can\nintroduce unnecessary costs and produce predictions that are not stable across\nindependent runs. To optimize this multi-criteria trade-off, we present a new\nframework that efficiently finds a principled balance between these three\nconsiderations using Bayesian optimisation. The pitfalls of optimising forest\nparameters purely for error reduction are demonstrated using two publicly\navailable real world datasets. We show that our framework leads to parameter\nsettings that are markedly different from the values discovered by error\nreduction metrics.","url_abs":"http://arxiv.org/abs/1706.09865v2","url_pdf":"http://arxiv.org/pdf/1706.09865v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalising-random-forest-parameter","repo_url":"https://github.com/liuchbryan/generalised_forest_tuning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"bayesian-optimisation","task_name":"Bayesian Optimisation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.09865","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}