{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/off-policy-evaluation-and-learning-for","title":"Off-Policy Evaluation and Learning for External Validity under a Covariate Shift","arxiv_id":"2002.11642","date":"2020-02-26","proceeding":"NeurIPS 2020 12","authors":["Masahiro Kato","Masatoshi Uehara","Shota Yasui"],"abstract":"We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new policy over the evaluation data, and that of off-policy learning (OPL) is to find a new policy that maximizes the expected reward over the evaluation data. Although the standard OPE and OPL assume the same distribution of covariate between the historical and evaluation data, a covariate shift often exists, i.e., the distribution of the covariate of the historical data is different from that of the evaluation data. In this paper, we derive the efficiency bound of OPE under a covariate shift. Then, we propose doubly robust and efficient estimators for OPE and OPL under a covariate shift by using a nonparametric estimator of the density ratio between the historical and evaluation data distributions. We also discuss other possible estimators and compare their theoretical properties. Finally, we confirm the effectiveness of the proposed estimators through experiments.","url_abs":"https://arxiv.org/abs/2002.11642v3","url_pdf":"https://arxiv.org/pdf/2002.11642v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"off-policy-evaluation-and-learning-for","repo_url":"https://github.com/MasaKat0/OPE_CS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"off-policy-evaluation","task_name":"Off-policy evaluation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2002.11642","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2002.11642"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MasaKat0/OPE_CS","reach":null}],"summary":{"ran_draft_wrong":2,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b56a62e42a86420c","entry":"process_args","repo":"MasaKat0/OPE_CS","repo_kind":"official","path":"cs_ope/experiments/experiment_evaluation.py","file_url":"https://github.com/MasaKat0/OPE_CS/blob/HEAD/cs_ope/experiments/experiment_evaluation.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b56a62e42a86420c"}},{"code_sha256_prefix":"06ebfec0a57daa95","entry":"process_args","repo":"MasaKat0/OPE_CS","repo_kind":"official","path":"cs_ope/experiments/experiment_learning.py","file_url":"https://github.com/MasaKat0/OPE_CS/blob/HEAD/cs_ope/experiments/experiment_learning.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"06ebfec0a57daa95"}},{"code_sha256_prefix":"390de638e884cd30","entry":"behavior_and_evaluation_policy","repo":"MasaKat0/OPE_CS","repo_kind":"official","path":"cs_ope/experiments/experiment_evaluation.py","file_url":"https://github.com/MasaKat0/OPE_CS/blob/HEAD/cs_ope/experiments/experiment_evaluation.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"390de638e884cd30"}},{"code_sha256_prefix":"15eac07dc38774a0","entry":"data_generation","repo":"MasaKat0/OPE_CS","repo_kind":"official","path":"cs_ope/experiments/experiment_evaluation.py","file_url":"https://github.com/MasaKat0/OPE_CS/blob/HEAD/cs_ope/experiments/experiment_evaluation.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"15eac07dc38774a0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}