{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/leveraging-factored-action-spaces-for-off","title":"Leveraging Factored Action Spaces for Off-Policy Evaluation","arxiv_id":"2307.07014","date":"2023-07-13","proceeding":null,"authors":["Aaman Rebello","Shengpu Tang","Jenna Wiens","Sonali Parbhoo"],"abstract":"Off-policy evaluation (OPE) aims to estimate the benefit of following a counterfactual sequence of actions, given data collected from executed sequences. However, existing OPE estimators often exhibit high bias and high variance in problems involving large, combinatorial action spaces. We investigate how to mitigate this issue using factored action spaces i.e. expressing each action as a combination of independent sub-actions from smaller action spaces. This approach facilitates a finer-grained analysis of how actions differ in their effects. In this work, we propose a new family of \"decomposed\" importance sampling (IS) estimators based on factored action spaces. Given certain assumptions on the underlying problem structure, we prove that the decomposed IS estimators have less variance than their original non-decomposed versions, while preserving the property of zero bias. Through simulations, we empirically verify our theoretical results, probing the validity of various assumptions. Provided with a technique that can derive the action space factorisation for a given problem, our work shows that OPE can be improved \"for free\" by utilising this inherent problem structure.","url_abs":"https://arxiv.org/abs/2307.07014v1","url_pdf":"https://arxiv.org/pdf/2307.07014v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"leveraging-factored-action-spaces-for-off","repo_url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"off-policy-evaluation","task_name":"Off-policy evaluation"},{"task_slug":null,"task_name":"counterfactual"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2307.07014","atlas_url":"https://app.syntology.ai/?focus=2307.07014","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.07014"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope","reach":null}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"4fefb8eb91cec738","entry":"off_policy_IS_estimator","repo":"ai4ai-lab/factored-action-spaces-for-ope","repo_kind":"official","path":"policy_estimators.py","file_url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope/blob/HEAD/policy_estimators.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4fefb8eb91cec738"}},{"code_sha256_prefix":"3632b55f7366702c","entry":"off_policy_PDIS_estimator","repo":"ai4ai-lab/factored-action-spaces-for-ope","repo_kind":"official","path":"policy_estimators.py","file_url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope/blob/HEAD/policy_estimators.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3632b55f7366702c"}},{"code_sha256_prefix":"dda6442b5c305f89","entry":"on_policy_Q_estimate","repo":"ai4ai-lab/factored-action-spaces-for-ope","repo_kind":"official","path":"policy_estimators.py","file_url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope/blob/HEAD/policy_estimators.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dda6442b5c305f89"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}