{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/leveraging-factored-action-spaces-for","title":"Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in Healthcare","arxiv_id":"2305.01738","date":"2023-05-02","proceeding":null,"authors":["Shengpu Tang","Maggie Makar","Michael W. Sjoding","Finale Doshi-Velez","Jenna Wiens"],"abstract":"Many reinforcement learning (RL) applications have combinatorial action spaces, where each action is a composition of sub-actions. A standard RL approach ignores this inherent factorization structure, resulting in a potential failure to make meaningful inferences about rarely observed sub-action combinations; this is particularly problematic for offline settings, where data may be limited. In this work, we propose a form of linear Q-function decomposition induced by factored action spaces. We study the theoretical properties of our approach, identifying scenarios where it is guaranteed to lead to zero bias when used to approximate the Q-function. Outside the regimes with theoretical guarantees, we show that our approach can still be useful because it leads to better sample efficiency without necessarily sacrificing policy optimality, allowing us to achieve a better bias-variance trade-off. Across several offline RL problems using simulators and real-world datasets motivated by healthcare, we demonstrate that incorporating factored action spaces into value-based RL can result in better-performing policies. Our approach can help an agent make more accurate inferences within underexplored regions of the state-action space when applying RL to observational datasets.","url_abs":"https://arxiv.org/abs/2305.01738v1","url_pdf":"https://arxiv.org/pdf/2305.01738v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"leveraging-factored-action-spaces-for","repo_url":"https://github.com/mld3/offlinerl_factoredactions","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"leveraging-factored-action-spaces-for","repo_url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"offline-rl","task_name":"Offline RL"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.01738","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.01738"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mld3/offlinerl_factoredactions","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope","reach":null}],"summary":{"ran_violates":1,"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1},"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"45ed9a311e055772","entry":"convert_factored_action","repo":"mld3/offlinerl_factoredactions","repo_kind":"official","path":"sepsisSim/exp-nets/run-NFQ_factored.py","file_url":"https://github.com/mld3/offlinerl_factoredactions/blob/HEAD/sepsisSim/exp-nets/run-NFQ_factored.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"45ed9a311e055772"}},{"code_sha256_prefix":"64e397c53381f27b","entry":"off_policy_DecIS_estimator","repo":"ai4ai-lab/factored-action-spaces-for-ope","repo_kind":"listed","path":"policy_estimators.py","file_url":"https://github.com/ai4ai-lab/factored-action-spaces-for-ope/blob/HEAD/policy_estimators.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"64e397c53381f27b"}},{"code_sha256_prefix":"15c42c6962dd8e88","entry":"init_networks","repo":"mld3/offlinerl_factoredactions","repo_kind":"official","path":"sepsisSim/exp-nets/run-NFQ_factored.py","file_url":"https://github.com/mld3/offlinerl_factoredactions/blob/HEAD/sepsisSim/exp-nets/run-NFQ_factored.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"15c42c6962dd8e88"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}