{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/missing-data-imputation-for-supervised","title":"Missing Data Imputation for Supervised Learning","arxiv_id":"1610.09075","date":"2016-10-28","proceeding":null,"authors":["Jason Poulos","Rafael Valle"],"abstract":"Missing data imputation can help improve the performance of prediction models\nin situations where missing data hide useful information. This paper compares\nmethods for imputing missing categorical data for supervised classification\ntasks. We experiment on two machine learning benchmark datasets with missing\ncategorical data, comparing classifiers trained on non-imputed (i.e., one-hot\nencoded) or imputed data with different levels of additional missing-data\nperturbation. We show imputation methods can increase predictive accuracy in\nthe presence of missing-data perturbation, which can actually improve\nprediction accuracy by regularizing the classifier. We achieve the\nstate-of-the-art on the Adult dataset with missing-data perturbation and\nk-nearest-neighbors (k-NN) imputation.","url_abs":"http://arxiv.org/abs/1610.09075v2","url_pdf":"http://arxiv.org/pdf/1610.09075v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"missing-data-imputation-for-supervised","repo_url":"https://github.com/rafaelvalle/MDI","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"imputation","task_name":"Imputation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/general-classification-on-cvr","task":"General Classification","dataset":"CVR","model":"Decision Trees","rank_in_archive_order":1,"of":1,"metrics":{"Test error":"0.027 ± 0.006"},"uses_additional_data":false},{"leaderboard":"/sota/imputation-on-adult-data-set","task":"Imputation","dataset":"Adult","model":"ANN","rank_in_archive_order":1,"of":1,"metrics":{"Test error":"0.144 ± 0.06"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1610.09075","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}