{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-comparison-of-methods-for-model-selection","title":"A comparison of methods for model selection when estimating individual treatment effects","arxiv_id":"1804.05146","date":"2018-04-14","proceeding":null,"authors":["Alejandro Schuler","Michael Baiocchi","Robert Tibshirani","Nigam Shah"],"abstract":"Practitioners in medicine, business, political science, and other fields are\nincreasingly aware that decisions should be personalized to each patient,\ncustomer, or voter. A given treatment (e.g. a drug or advertisement) should be\nadministered only to those who will respond most positively, and certainly not\nto those who will be harmed by it. Individual-level treatment effects can be\nestimated with tools adapted from machine learning, but different models can\nyield contradictory estimates. Unlike risk prediction models, however,\ntreatment effect models cannot be easily evaluated against each other using a\nheld-out test set because the true treatment effect itself is never directly\nobserved. Besides outcome prediction accuracy, several metrics that can\nleverage held-out data to evaluate treatment effects models have been proposed,\nbut they are not widely used. We provide a didactic framework that elucidates\nthe relationships between the different approaches and compare them all using a\nvariety of simulations of both randomized and observational data. Our results\nshow that researchers estimating heterogenous treatment effects need not limit\nthemselves to a single model-fitting algorithm. Instead of relying on a single\nmethod, multiple models fit by a diverse set of algorithms should be evaluated\nagainst each other using an objective function learned from the validation set.\nThe model minimizing that objective should be used for estimating the\nindividual treatment effect for future individuals.","url_abs":"http://arxiv.org/abs/1804.05146v2","url_pdf":"http://arxiv.org/pdf/1804.05146v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-comparison-of-methods-for-model-selection","repo_url":"https://github.com/som-shahlab/ITE-model-selection","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"a-comparison-of-methods-for-model-selection","repo_url":"https://github.com/Ibotta/mr_uplift","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"a-comparison-of-methods-for-model-selection","repo_url":"https://github.com/tonyduan/hte-prediction-rcts","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"model-selection","task_name":"Model Selection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.05146","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1804.05146"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tonyduan/hte-prediction-rcts","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Ibotta/mr_uplift","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/som-shahlab/ITE-model-selection","reach":{"status":"ok"}}],"summary":{"unverified":6},"by_repo_kind":{"listed":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7158f0b5c2fbe552","entry":"bootstrap_dataset","repo":"tonyduan/hte-prediction-rcts","repo_kind":"listed","path":"src/dataloader.py","file_url":"https://github.com/tonyduan/hte-prediction-rcts/blob/HEAD/src/dataloader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7158f0b5c2fbe552"}},{"code_sha256_prefix":"db34a7d767a49b30","entry":"bucket_arr","repo":"tonyduan/hte-prediction-rcts","repo_kind":"listed","path":"src/evaluate.py","file_url":"https://github.com/tonyduan/hte-prediction-rcts/blob/HEAD/src/evaluate.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"db34a7d767a49b30"}},{"code_sha256_prefix":"1dcdabe69815a74e","entry":"get_range","repo":"tonyduan/hte-prediction-rcts","repo_kind":"listed","path":"src/evaluate.py","file_url":"https://github.com/tonyduan/hte-prediction-rcts/blob/HEAD/src/evaluate.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1dcdabe69815a74e"}},{"code_sha256_prefix":"80132d634ab80ed9","entry":"load_data","repo":"tonyduan/hte-prediction-rcts","repo_kind":"listed","path":"src/dataloader.py","file_url":"https://github.com/tonyduan/hte-prediction-rcts/blob/HEAD/src/dataloader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"80132d634ab80ed9"}},{"code_sha256_prefix":"b5e9d85bbda9441b","entry":"run_for_ascvd","repo":"tonyduan/hte-prediction-rcts","repo_kind":"listed","path":"src/baselines.py","file_url":"https://github.com/tonyduan/hte-prediction-rcts/blob/HEAD/src/baselines.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b5e9d85bbda9441b"}},{"code_sha256_prefix":"ff7043251da03ae4","entry":"wald_test","repo":"tonyduan/hte-prediction-rcts","repo_kind":"listed","path":"src/evaluate.py","file_url":"https://github.com/tonyduan/hte-prediction-rcts/blob/HEAD/src/evaluate.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ff7043251da03ae4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}