{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/quantile-constrained-reinforcement-learning-a","title":"Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability","arxiv_id":"2211.15034","date":"2022-11-28","proceeding":null,"authors":["Whiyoung Jung","Myungsik Cho","Jongeui Park","Youngchul Sung"],"abstract":"Constrained reinforcement learning (RL) is an area of RL whose objective is to find an optimal policy that maximizes expected cumulative return while satisfying a given constraint. Most of the previous constrained RL works consider expected cumulative sum cost as the constraint. However, optimization with this constraint cannot guarantee a target probability of outage event that the cumulative sum cost exceeds a given threshold. This paper proposes a framework, named Quantile Constrained RL (QCRL), to constrain the quantile of the distribution of the cumulative sum cost that is a necessary and sufficient condition to satisfy the outage constraint. This is the first work that tackles the issue of applying the policy gradient theorem to the quantile and provides theoretical results for approximating the gradient of the quantile. Based on the derived theoretical results and the technique of the Lagrange multiplier, we construct a constrained RL algorithm named Quantile Constrained Policy Optimization (QCPO). We use distributional RL with the Large Deviation Principle (LDP) to estimate quantiles and tail probability of the cumulative sum cost for the implementation of QCPO. The implemented algorithm satisfies the outage probability constraint after the training period.","url_abs":"https://arxiv.org/abs/2211.15034v1","url_pdf":"https://arxiv.org/pdf/2211.15034v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"quantile-constrained-reinforcement-learning-a","repo_url":"https://github.com/wyjung0625/qcpo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2211.15034","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2211.15034"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/wyjung0625/QCPO","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wyjung0625/qcpo","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"57c952fdc1ca4a15","entry":"infill_info","repo":"wyjung0625/QCPO","repo_kind":"official","path":"qcpo/safety_gym_envs/safety_gym_env.py","file_url":"https://github.com/wyjung0625/QCPO/blob/HEAD/qcpo/safety_gym_envs/safety_gym_env.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"57c952fdc1ca4a15"}},{"code_sha256_prefix":"08f661b3dd76de88","entry":"normalize","repo":"wyjung0625/QCPO","repo_kind":"official","path":"qcpo/dist_rl_utils.py","file_url":"https://github.com/wyjung0625/QCPO/blob/HEAD/qcpo/dist_rl_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"08f661b3dd76de88"}},{"code_sha256_prefix":"c3eedd504cc05a91","entry":"quantile_huber_loss","repo":"wyjung0625/QCPO","repo_kind":"official","path":"qcpo/dist_rl_utils.py","file_url":"https://github.com/wyjung0625/QCPO/blob/HEAD/qcpo/dist_rl_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c3eedd504cc05a91"}},{"code_sha256_prefix":"7ed1ec4978aa12f1","entry":"weibull_tail_loss","repo":"wyjung0625/QCPO","repo_kind":"official","path":"qcpo/dist_rl_utils.py","file_url":"https://github.com/wyjung0625/QCPO/blob/HEAD/qcpo/dist_rl_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7ed1ec4978aa12f1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}