{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-consistency-of-supervised-learning","title":"On the consistency of supervised learning with missing values","arxiv_id":"1902.06931","date":"2019-02-19","proceeding":null,"authors":["Julie Josse","Jacob M. Chen","Nicolas Prost","Erwan Scornet","Gaël Varoquaux"],"abstract":"In many application settings, the data have missing entries which make analysis challenging. An abundant literature addresses missing values in an inferential framework: estimating parameters and their variance from incomplete tables. Here, we consider supervised-learning settings: predicting a target when missing values appear in both training and testing data. We show the consistency of two approaches in prediction. A striking result is that the widely-used method of imputing with a constant, such as the mean prior to learning is consistent when missing values are not informative. This contrasts with inferential settings where mean imputation is pointed at for distorting the distribution of the data. That such a simple approach can be consistent is important in practice. We also show that a predictor suited for complete observations can predict optimally on incomplete data, through multiple imputation. Finally, to compare imputation with learning directly with a model that accounts for missing values, we analyze further decision trees. These can naturally tackle empirical risk minimization with missing values, due to their ability to handle the half-discrete nature of incomplete variables. After comparing theoretically and empirically different missing values strategies in trees, we recommend using the \"missing incorporated in attribute\" method as it can handle both non-informative and informative missing values.","url_abs":"https://arxiv.org/abs/1902.06931v5","url_pdf":"https://arxiv.org/pdf/1902.06931v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-consistency-of-supervised-learning","repo_url":"https://github.com/dirty-data/supervised_missing","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"on-the-consistency-of-supervised-learning","repo_url":"https://github.com/jacobmchen/supervised_missing","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"on-the-consistency-of-supervised-learning","repo_url":"https://github.com/nprost/supervised_missing","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"imputation","task_name":"Imputation"},{"task_slug":"missing-values","task_name":"Missing Values"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1902.06931","atlas_url":"https://app.syntology.ai/?focus=1902.06931","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1902.06931"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nprost/supervised_missing","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jacobmchen/supervised_missing","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dirty-data/supervised_missing","reach":{"status":"ok"}}],"summary":{"ran_violates":2,"ran_honours":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"e09fe9ef3911744a","entry":"critmia","repo":"nprost/supervised_missing","repo_kind":"official","path":"analysis/computation_theoretical_risk.py","file_url":"https://github.com/nprost/supervised_missing/blob/HEAD/analysis/computation_theoretical_risk.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e09fe9ef3911744a"}},{"code_sha256_prefix":"44f83197999b0597","entry":"riskblock","repo":"nprost/supervised_missing","repo_kind":"official","path":"analysis/computation_theoretical_risk.py","file_url":"https://github.com/nprost/supervised_missing/blob/HEAD/analysis/computation_theoretical_risk.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"44f83197999b0597"}},{"code_sha256_prefix":"c1bd75be1ce2d43b","entry":"riskmia","repo":"nprost/supervised_missing","repo_kind":"official","path":"analysis/computation_theoretical_risk.py","file_url":"https://github.com/nprost/supervised_missing/blob/HEAD/analysis/computation_theoretical_risk.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c1bd75be1ce2d43b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}