{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/conservative-offline-distributional","title":"Conservative Offline Distributional Reinforcement Learning","arxiv_id":"2107.06106","date":"2021-07-12","proceeding":"NeurIPS 2021 12","authors":["Yecheng Jason Ma","Dinesh Jayaraman","Osbert Bastani"],"abstract":"Many reinforcement learning (RL) problems in practice are offline, learning purely from observational data. A key challenge is how to ensure the learned policy is safe, which requires quantifying the risk associated with different actions. In the online setting, distributional RL algorithms do so by learning the distribution over returns (i.e., cumulative rewards) instead of the expected return; beyond quantifying risk, they have also been shown to learn better representations for planning. We propose Conservative Offline Distributional Actor Critic (CODAC), an offline RL algorithm suitable for both risk-neutral and risk-averse domains. CODAC adapts distributional RL to the offline setting by penalizing the predicted quantiles of the return for out-of-distribution actions. We prove that CODAC learns a conservative return distribution -- in particular, for finite MDPs, CODAC converges to an uniform lower bound on the quantiles of the return distribution; our proof relies on a novel analysis of the distributional Bellman operator. In our experiments, on two challenging robot navigation tasks, CODAC successfully learns risk-averse policies using offline data collected purely from risk-neutral agents. Furthermore, CODAC is state-of-the-art on the D4RL MuJoCo benchmark in terms of both expected and risk-sensitive performance.","url_abs":"https://arxiv.org/abs/2107.06106v2","url_pdf":"https://arxiv.org/pdf/2107.06106v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"conservative-offline-distributional","repo_url":"https://github.com/JasonMa2016/CODAC","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"d4rl","task_name":"D4RL"},{"task_slug":"distributional-reinforcement-learning","task_name":"Distributional Reinforcement Learning"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"offline-rl","task_name":"Offline RL"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"robot-navigation","task_name":"Robot Navigation"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2107.06106","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2107.06106"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/JasonMa2016/CODAC","reach":null}],"summary":{"ran":3,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"15db1b8d1bff5a1e","entry":"GaussianPolicy","repo":"JasonMa2016/CODAC","repo_kind":"official","path":"distributional/codac.py","file_url":"https://github.com/JasonMa2016/CODAC/blob/HEAD/distributional/codac.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"15db1b8d1bff5a1e"}},{"code_sha256_prefix":"04449fb6e1def94c","entry":"LinearSchedule","repo":"JasonMa2016/CODAC","repo_kind":"official","path":"distributional/codac.py","file_url":"https://github.com/JasonMa2016/CODAC/blob/HEAD/distributional/codac.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"04449fb6e1def94c"}},{"code_sha256_prefix":"47e205f9e3748c01","entry":"QuantileMlp","repo":"JasonMa2016/CODAC","repo_kind":"official","path":"distributional/codac.py","file_url":"https://github.com/JasonMa2016/CODAC/blob/HEAD/distributional/codac.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"47e205f9e3748c01"}},{"code_sha256_prefix":"ca7633e5aeaae825","entry":"quantile_regression_loss","repo":"JasonMa2016/CODAC","repo_kind":"official","path":"distributional/codac.py","file_url":"https://github.com/JasonMa2016/CODAC/blob/HEAD/distributional/codac.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ca7633e5aeaae825"}},{"code_sha256_prefix":"e6a33156db616c93","entry":"CODAC","repo":"JasonMa2016/CODAC","repo_kind":"official","path":"distributional/codac.py","file_url":"https://github.com/JasonMa2016/CODAC/blob/HEAD/distributional/codac.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e6a33156db616c93"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}