{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/feasible-actor-critic-constrained","title":"Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety","arxiv_id":"2105.10682","date":"2021-05-22","proceeding":null,"authors":["Haitong Ma","Yang Guan","Shegnbo Eben Li","Xiangteng Zhang","Sifa Zheng","Jianyu Chen"],"abstract":"The safety constraints commonly used by existing safe reinforcement learning (RL) methods are defined only on expectation of initial states, but allow each certain state to be unsafe, which is unsatisfying for real-world safety-critical tasks. In this paper, we introduce the feasible actor-critic (FAC) algorithm, which is the first model-free constrained RL method that considers statewise safety, e.g, safety for each initial state. We claim that some states are inherently unsafe no matter what policy we choose, while for other states there exist policies ensuring safety, where we say such states and policies are feasible. By constructing a statewise Lagrange function available on RL sampling and adopting an additional neural network to approximate the statewise Lagrange multiplier, we manage to obtain the optimal feasible policy which ensures safety for each feasible state and the safest possible policy for infeasible states. Furthermore, the trained multiplier net can indicate whether a given state is feasible or not through the statewise complementary slackness condition. We provide theoretical guarantees that FAC outperforms previous expectation-based constrained RL methods in terms of both constraint satisfaction and reward optimization. Experimental results on both robot locomotive tasks and safe exploration tasks verify the safety enhancement and feasibility interpretation of the proposed method.","url_abs":"https://arxiv.org/abs/2105.10682v3","url_pdf":"https://arxiv.org/pdf/2105.10682v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"feasible-actor-critic-constrained","repo_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"feasible-actor-critic-constrained","repo_url":"https://github.com/mahaitongdae/safety_index_synthesis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"feasible-actor-critic-constrained","repo_url":"https://github.com/zlr20/saferl_kit","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-exploration","task_name":"Safe Exploration"},{"task_slug":"safe-reinforcement-learning","task_name":"Safe Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2105.10682","atlas_url":"https://app.syntology.ai/?focus=2105.10682","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2105.10682"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mahaitongdae/safety_index_synthesis","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mahaitongdae/Feasible-Actor-Critic","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zlr20/saferl_kit","reach":null}],"summary":{"ran_honours":1,"unverified":5},"by_repo_kind":{"official":{"samples":6,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3ebb420a5d10cc64","entry":"update_mean_var_count_from_moments","repo":"mahaitongdae/Feasible-Actor-Critic","repo_kind":"official","path":"preprocessor.py","file_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic/blob/HEAD/preprocessor.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3ebb420a5d10cc64"}},{"code_sha256_prefix":"8e77d175d25461f7","entry":"built_parser","repo":"mahaitongdae/Feasible-Actor-Critic","repo_kind":"official","path":"train_script4fsac.py","file_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic/blob/HEAD/train_script4fsac.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8e77d175d25461f7"}},{"code_sha256_prefix":"7fb1719dd18effc4","entry":"built_parser","repo":"mahaitongdae/Feasible-Actor-Critic","repo_kind":"official","path":"train_scripts/train_script4fsac_with_trained_model.py","file_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic/blob/HEAD/train_scripts/train_script4fsac_with_trained_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7fb1719dd18effc4"}},{"code_sha256_prefix":"0b9cc3b9817b83d6","entry":"compute_convergence_speed","repo":"mahaitongdae/Feasible-Actor-Critic","repo_kind":"official","path":"ploter.py","file_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic/blob/HEAD/ploter.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0b9cc3b9817b83d6"}},{"code_sha256_prefix":"6d7182a8d3f335dd","entry":"help_func","repo":"mahaitongdae/Feasible-Actor-Critic","repo_kind":"official","path":"ploter.py","file_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic/blob/HEAD/ploter.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6d7182a8d3f335dd"}},{"code_sha256_prefix":"ce7191c627d919b5","entry":"min_n","repo":"mahaitongdae/Feasible-Actor-Critic","repo_kind":"official","path":"ploter.py","file_url":"https://github.com/mahaitongdae/Feasible-Actor-Critic/blob/HEAD/ploter.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ce7191c627d919b5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}