{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-trust-region-based-safe","title":"Trust Region-Based Safe Distributional Reinforcement Learning for Multiple Constraints","arxiv_id":"2301.10923","date":"2023-01-26","proceeding":"NeurIPS 2023 11","authors":["Dohyeong Kim","Kyungjae Lee","Songhwai Oh"],"abstract":"In safety-critical robotic tasks, potential failures must be reduced, and multiple constraints must be met, such as avoiding collisions, limiting energy consumption, and maintaining balance. Thus, applying safe reinforcement learning (RL) in such robotic tasks requires to handle multiple constraints and use risk-averse constraints rather than risk-neutral constraints. To this end, we propose a trust region-based safe RL algorithm for multiple constraints called a safe distributional actor-critic (SDAC). Our main contributions are as follows: 1) introducing a gradient integration method to manage infeasibility issues in multi-constrained problems, ensuring theoretical convergence, and 2) developing a TD($\\lambda$) target distribution to estimate risk-averse constraints with low biases. We evaluate SDAC through extensive experiments involving multi- and single-constrained robotic tasks. While maintaining high scores, SDAC shows 1.93 times fewer steps to satisfy all constraints in multi-constrained tasks and 1.78 times fewer constraint violations in single-constrained tasks compared to safe RL baselines. Code is available at: https://github.com/rllab-snu/Safe-Distributional-Actor-Critic.","url_abs":"https://arxiv.org/abs/2301.10923v2","url_pdf":"https://arxiv.org/pdf/2301.10923v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-trust-region-based-safe","repo_url":"https://github.com/rllab-snu/safe-distributional-actor-critic","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"distributional-reinforcement-learning","task_name":"Distributional Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2301.10923","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2301.10923"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/rllab-snu/Safe-Distributional-Actor-Critic","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rllab-snu/safe-distributional-actor-critic","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7a753817cb0801b7","entry":"bt","repo":"rllab-snu/Safe-Distributional-Actor-Critic","repo_kind":"official","path":"safety_gym/cvpo/safe_rl/policy/cvpo.py","file_url":"https://github.com/rllab-snu/Safe-Distributional-Actor-Critic/blob/HEAD/safety_gym/cvpo/safe_rl/policy/cvpo.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7a753817cb0801b7"}},{"code_sha256_prefix":"548253ae95d04e97","entry":"btr","repo":"rllab-snu/Safe-Distributional-Actor-Critic","repo_kind":"official","path":"safety_gym/cvpo/safe_rl/policy/cvpo.py","file_url":"https://github.com/rllab-snu/Safe-Distributional-Actor-Critic/blob/HEAD/safety_gym/cvpo/safe_rl/policy/cvpo.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"548253ae95d04e97"}},{"code_sha256_prefix":"dff135af934deab2","entry":"safe_inverse","repo":"rllab-snu/Safe-Distributional-Actor-Critic","repo_kind":"official","path":"safety_gym/cvpo/safe_rl/policy/cvpo.py","file_url":"https://github.com/rllab-snu/Safe-Distributional-Actor-Critic/blob/HEAD/safety_gym/cvpo/safe_rl/policy/cvpo.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dff135af934deab2"}},{"code_sha256_prefix":"2bc10041d414773a","entry":"mlp","repo":"rllab-snu/Safe-Distributional-Actor-Critic","repo_kind":"official","path":"safety_gym/cvpo/safe_rl/policy/model/mlp_ac.py","file_url":"https://github.com/rllab-snu/Safe-Distributional-Actor-Critic/blob/HEAD/safety_gym/cvpo/safe_rl/policy/model/mlp_ac.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2bc10041d414773a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}