{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/constrained-policy-optimization","title":"Constrained Policy Optimization","arxiv_id":"1705.10528","date":"2017-05-30","proceeding":"ICML 2017 8","authors":["Joshua Achiam","David Held","Aviv Tamar","Pieter Abbeel"],"abstract":"For many applications of reinforcement learning it can be more convenient to\nspecify both a reward function and constraints, rather than trying to design\nbehavior through the reward function. For example, systems that physically\ninteract with or around humans should satisfy safety constraints. Recent\nadvances in policy search algorithms (Mnih et al., 2016, Schulman et al., 2015,\nLillicrap et al., 2016, Levine et al., 2016) have enabled new capabilities in\nhigh-dimensional control, but do not consider the constrained setting.\n  We propose Constrained Policy Optimization (CPO), the first general-purpose\npolicy search algorithm for constrained reinforcement learning with guarantees\nfor near-constraint satisfaction at each iteration. Our method allows us to\ntrain neural network policies for high-dimensional control while making\nguarantees about policy behavior all throughout training. Our guarantees are\nbased on a new theoretical result, which is of independent interest: we prove a\nbound relating the expected returns of two policies to an average divergence\nbetween them. We demonstrate the effectiveness of our approach on simulated\nrobot locomotion tasks where the agent must satisfy constraints motivated by\nsafety.","url_abs":"http://arxiv.org/abs/1705.10528v1","url_pdf":"http://arxiv.org/pdf/1705.10528v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/awesomericky/sim2real-grounding-simulator","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/dobro12/CPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/ethz-asl/rl-navigation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/hari-sikchi/pytorch_CPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/jachiam/cpo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/jemaw/gym-safety","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/sapanachaudhary/pytorch-cpo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/varunjain3/SafetyRL_HighwayEnv","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"constrained-policy-optimization","repo_url":"https://github.com/ymzhang01/mujoco-circle","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.10528","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1705.10528"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jachiam/cpo","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ymzhang01/mujoco-circle","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/awesomericky/sim2real-grounding-simulator","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sapanachaudhary/pytorch-cpo","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/varunjain3/SafetyRL_HighwayEnv","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jemaw/gym-safety","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ethz-asl/rl-navigation","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dobro12/CPO","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hari-sikchi/pytorch_CPO","reach":{"status":"unanswered"}}],"summary":{"ran_honours":2,"ran_draft_wrong":1,"unverified":3},"by_repo_kind":{"listed":{"samples":6,"ran":3,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"b50dc1dd1b07f825","entry":"flatten","repo":"Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO","repo_kind":"listed","path":"CPO.py","file_url":"https://github.com/Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO/blob/HEAD/CPO.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":2,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b50dc1dd1b07f825"}},{"code_sha256_prefix":"f0d38d6259a00367","entry":"get_flat_params","repo":"Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO","repo_kind":"listed","path":"CPO.py","file_url":"https://github.com/Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO/blob/HEAD/CPO.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f0d38d6259a00367"}},{"code_sha256_prefix":"8ef3c4473d33a01b","entry":"line_search","repo":"Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO","repo_kind":"listed","path":"CPO.py","file_url":"https://github.com/Bigpig4396/PyTorch-Constrained-Policy-Optimization-CPO/blob/HEAD/CPO.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8ef3c4473d33a01b"}},{"code_sha256_prefix":"aad2ffc14dc1c396","entry":"conjugate_gradients","repo":"sapanachaudhary/pytorch-cpo","repo_kind":"listed","path":"algos/cpo.py","file_url":"https://github.com/sapanachaudhary/pytorch-cpo/blob/HEAD/algos/cpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"aad2ffc14dc1c396"}},{"code_sha256_prefix":"0feb3a0151752db0","entry":"cpo_step","repo":"sapanachaudhary/pytorch-cpo","repo_kind":"listed","path":"algos/cpo.py","file_url":"https://github.com/sapanachaudhary/pytorch-cpo/blob/HEAD/algos/cpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0feb3a0151752db0"}},{"code_sha256_prefix":"f7c68d288ffc351b","entry":"line_search","repo":"sapanachaudhary/pytorch-cpo","repo_kind":"listed","path":"algos/cpo.py","file_url":"https://github.com/sapanachaudhary/pytorch-cpo/blob/HEAD/algos/cpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f7c68d288ffc351b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}