{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/training-big-random-forests-with-little","title":"Training Big Random Forests with Little Resources","arxiv_id":"1802.06394","date":"2018-02-18","proceeding":null,"authors":["Fabian Gieseke","Christian Igel"],"abstract":"Without access to large compute clusters, building random forests on large\ndatasets is still a challenging problem. This is, in particular, the case if\nfully-grown trees are desired. We propose a simple yet effective framework that\nallows to efficiently construct ensembles of huge trees for hundreds of\nmillions or even billions of training instances using a cheap desktop computer\nwith commodity hardware. The basic idea is to consider a multi-level\nconstruction scheme, which builds top trees for small random subsets of the\navailable data and which subsequently distributes all training instances to the\ntop trees' leaves for further processing. While being conceptually simple, the\noverall efficiency crucially depends on the particular implementation of the\ndifferent phases. The practical merits of our approach are demonstrated using\ndense datasets with hundreds of millions of training instances.","url_abs":"http://arxiv.org/abs/1802.06394v1","url_pdf":"http://arxiv.org/pdf/1802.06394v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"training-big-random-forests-with-little","repo_url":"https://github.com/gieseke/woody","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}