{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/optimistic-distributionally-robust-policy","title":"Optimistic Distributionally Robust Policy Optimization","arxiv_id":"2006.07815","date":"2020-06-14","proceeding":null,"authors":["Jun Song","Chaoyue Zhao"],"abstract":"Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), as the widely employed policy based reinforcement learning (RL) methods, are prone to converge to a sub-optimal solution as they limit the policy representation to a particular parametric distribution class. To address this issue, we develop an innovative Optimistic Distributionally Robust Policy Optimization (ODRPO) algorithm, which effectively utilizes Optimistic Distributionally Robust Optimization (DRO) approach to solve the trust region constrained optimization problem without parameterizing the policies. Our algorithm improves TRPO and PPO with a higher sample efficiency and a better performance of the final policy while attaining the learning stability. Moreover, it achieves a globally optimal policy update that is not promised in the prevailing policy based RL algorithms. Experiments across tabular domains and robotic locomotion tasks demonstrate the effectiveness of our approach.","url_abs":"https://arxiv.org/abs/2006.07815v1","url_pdf":"https://arxiv.org/pdf/2006.07815v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"optimistic-distributionally-robust-policy","repo_url":"https://github.com/kadysongbb/dr-trpo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"entropy-regularization","method_name":"Entropy Regularization"},{"method_slug":"ppo","method_name":"PPO"},{"method_slug":"trpo","method_name":"TRPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2006.07815","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2006.07815"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kadysongbb/dr-trpo","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"00277beee4c6e0a9","entry":"denormalize","repo":"kadysongbb/dr-trpo","repo_kind":"official","path":"continuous_control/GAC/helpers.py","file_url":"https://github.com/kadysongbb/dr-trpo/blob/HEAD/continuous_control/GAC/helpers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"00277beee4c6e0a9"}},{"code_sha256_prefix":"2e7a471dea032a49","entry":"episode_stats","repo":"kadysongbb/dr-trpo","repo_kind":"official","path":"tabular/train_helper.py","file_url":"https://github.com/kadysongbb/dr-trpo/blob/HEAD/tabular/train_helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2e7a471dea032a49"}},{"code_sha256_prefix":"4a935bce837bbcad","entry":"normalize","repo":"kadysongbb/dr-trpo","repo_kind":"official","path":"continuous_control/GAC/helpers.py","file_url":"https://github.com/kadysongbb/dr-trpo/blob/HEAD/continuous_control/GAC/helpers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4a935bce837bbcad"}},{"code_sha256_prefix":"376d051ff6f4515b","entry":"run_episode","repo":"kadysongbb/dr-trpo","repo_kind":"official","path":"tabular/train_helper.py","file_url":"https://github.com/kadysongbb/dr-trpo/blob/HEAD/tabular/train_helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"376d051ff6f4515b"}},{"code_sha256_prefix":"e4f0e8c345e82dae","entry":"run_policy","repo":"kadysongbb/dr-trpo","repo_kind":"official","path":"tabular/train_helper.py","file_url":"https://github.com/kadysongbb/dr-trpo/blob/HEAD/tabular/train_helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e4f0e8c345e82dae"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}