{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2509-26429","title":"An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes","arxiv_id":"2509.26429","date":"2025-09-30","proceeding":"ICLR","authors":["Emil Javurek","Valentyn Melnychuk","Jonas Schweisthal","Konstantin Hess","Dennis Frauen","Stefan Feuerriegel"],"abstract":"Predicting individualized potential outcomes in sequential decision-making is central for optimizing therapeutic decisions in personalized medicine (e.g., which dosing sequence to give to a cancer patient). However, predicting potential outcomes over long horizons is notoriously difficult. Existing methods that break the curse of the horizon typically lack strong theoretical guarantees such as orthogonality and quasi-oracle efficiency. In this paper, we revisit the problem of predicting individualized potential outcomes in sequential decision-making (i.e., estimating Q-functions in Markov decision processes with observational data) through a causal inference lens. In particular, we develop a comprehensive theoretical foundation for meta-learners in this setting with a focus on beneficial theoretical properties. As a result, we yield a novel meta-learner called DRQ-learner and establish that it is: (1) doubly robust (i.e., valid inference under the misspecification of one of the nuisances), (2) Neyman-orthogonal (i.e., insensitive to first-order estimation errors in the nuisance functions), and (3) achieves quasi-oracle efficiency (i.e., behaves asymptotically as if the ground-truth nuisance functions were known). Our DRQ-learner is applicable to settings with both discrete and continuous state spaces. Further, our DRQ-learner is flexible and can be used together with arbitrary machine learning models (e.g., neural networks). We validate our theoretical results through numerical experiments, thereby showing that our meta-learner outperforms state-of-the-art baselines.","url_abs":"https://arxiv.org/abs/2509.26429","url_pdf":"https://arxiv.org/pdf/2509.26429","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2509.26429","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2509.26429"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs","reach":{"status":"ok"}}],"summary":{"unverified":6},"by_repo_kind":{"found_in_text":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"bbe2abe344b5ef99","entry":"get","repo":"EmilJavurek/Orthogonal-Q-in-MDPs","repo_kind":"found_in_text","path":"algos/registry.py","file_url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs/blob/HEAD/algos/registry.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bbe2abe344b5ef99"}},{"code_sha256_prefix":"58ee1038f5f50b19","entry":"online_rollout","repo":"EmilJavurek/Orthogonal-Q-in-MDPs","repo_kind":"found_in_text","path":"algos/online_prediction.py","file_url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs/blob/HEAD/algos/online_prediction.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"58ee1038f5f50b19"}},{"code_sha256_prefix":"37994848455847b9","entry":"reconstruct_trajectories","repo":"EmilJavurek/Orthogonal-Q-in-MDPs","repo_kind":"found_in_text","path":"algos/bench_Qreg.py","file_url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs/blob/HEAD/algos/bench_Qreg.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"37994848455847b9"}},{"code_sha256_prefix":"7035de5c053aa2b2","entry":"reconstruct_trajectories","repo":"EmilJavurek/Orthogonal-Q-in-MDPs","repo_kind":"found_in_text","path":"algos/nuisance_densities.py","file_url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs/blob/HEAD/algos/nuisance_densities.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7035de5c053aa2b2"}},{"code_sha256_prefix":"b7ca1a731038ba64","entry":"register","repo":"EmilJavurek/Orthogonal-Q-in-MDPs","repo_kind":"found_in_text","path":"algos/registry.py","file_url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs/blob/HEAD/algos/registry.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b7ca1a731038ba64"}},{"code_sha256_prefix":"09534533a2617e04","entry":"torch_atleast_2d","repo":"EmilJavurek/Orthogonal-Q-in-MDPs","repo_kind":"found_in_text","path":"algos/online_prediction.py","file_url":"https://github.com/EmilJavurek/Orthogonal-Q-in-MDPs/blob/HEAD/algos/online_prediction.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"09534533a2617e04"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"stat.ML","source":"arxiv_api"},"syntology_extracted_results":null}