{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/who-wins-the-miss-contest-for-imputation","title":"Who wins the Miss Contest for Imputation Methods? Our Vote for Miss BooPF","arxiv_id":"1711.11394","date":"2017-11-30","proceeding":null,"authors":["Burim Ramosaj","Markus Pauly"],"abstract":"Missing data is an expected issue when large amounts of data is collected,\nand several imputation techniques have been proposed to tackle this problem.\nBeneath classical approaches such as MICE, the application of Machine Learning\ntechniques is tempting. Here, the recently proposed missForest imputation\nmethod has shown high imputation accuracy under the Missing (Completely) at\nRandom scheme with various missing rates. In its core, it is based on a random\nforest for classification and regression, respectively. In this paper we study\nwhether this approach can even be enhanced by other methods such as the\nstochastic gradient tree boosting method, the C5.0 algorithm or modified random\nforest procedures. In particular, other resampling strategies within the random\nforest protocol are suggested. In an extensive simulation study, we analyze\ntheir performances for continuous, categorical as well as mixed-type data.\nTherein, MissBooPF, a combination of the stochastic gradient tree boosting\nmethod together with the parametrically bootstrapped random forest method,\nappeared to be promising. Finally, an empirical analysis focusing on credit\ninformation and Facebook data is conducted.","url_abs":"http://arxiv.org/abs/1711.11394v1","url_pdf":"http://arxiv.org/pdf/1711.11394v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"who-wins-the-miss-contest-for-imputation","repo_url":"https://github.com/yaroslav-moiseev/evidence-based-possibly-best-practices-in-classical-ML","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"imputation","task_name":"Imputation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}