Browse State-of-the-Art › Safe Reinforcement Learning
Safe Reinforcement Learning
100 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 100 papers with code (306 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 May 2017 10 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 3 pointer-only (licence)For many applications of reinforcement learning it can be more convenient to specify both a reward function and constraints, rather than trying to design behavior through the reward function.
-
13 Aug 2021 5 repositories listedThe last half-decade has seen a steep rise in the number of contributions on safe learning methods for real-world robotic deployments from both the control and reinforcement learning communities.
-
2 May 2024 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Ensuring the safety of Reinforcement Learning (RL) is crucial for its deployment in real-world applications.
-
15 Jun 2023 3 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedThis paper presents a comprehensive benchmarking suite tailored to offline safe reinforcement learning (RL) challenges, aiming to foster progress in the development and evaluation of safe learning algorithms in both the…
-
15 Sep 2022 3 repositories listedCompared to previous safe RL methods, CUP enjoys the benefits of 1) CUP generalizes the surrogate functions to generalized advantage estimator (GAE), leading to strong empirical performance.
-
22 May 2021 3 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedThe safety constraints commonly used by existing safe reinforcement learning (RL) methods are defined only on expectation of initial states, but allow each certain state to be unsafe, which is unsatisfying for…
-
26 Jan 2024 2 repositories listed Syntology ran 13 of 22 samples · 9 unverified · 5 pointer-only (licence)Results on benchmark tasks show that our method not only achieves an asymptotic performance comparable to state-of-the-art on-policy methods while using much fewer samples, but also significantly reduces constraint…
-
23 Jan 2024 2 repositories listedReinforcement learning (RL) excels in applications such as video games, but ensuring safety as well as the ability to achieve the specified goals remains challenging when using RL for real-world problems, such as…
-
29 Oct 2022 2 repositories listedIn this work, we propose a self-improving artificial intelligence system to enhance the safety performance of reinforcement learning (RL)-based autonomous driving (AD) agents using black-box verification methods.
-
21 Jul 2022 2 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedWe introduce a general approach for seeking a stationary point in high dimensional non-linear stochastic optimization problems in which maintaining safety during learning is crucial.
-
16 May 2022 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 1 pointer-only (licence)Recent studies incorporate feasible sets into CRL with energy-based methods such as control barrier function (CBF), safety index (SI), and leverage prior conservative estimations of feasible sets, which harms the…
-
28 Jan 2022 2 repositories listed Syntology ran 5 of 11 samples · 6 unverified · 6 pointer-only (licence)Safe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications.
-
26 Sep 2021 2 repositories listedBased on MetaDrive, we construct a variety of RL tasks and baselines in both single-agent and multi-agent settings, including benchmarking generalizability across unseen scenes, safe exploration, and learning…
-
29 Oct 2020 2 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedSafety remains a central obstacle preventing widespread use of RL in the real world: learning new tasks in uncertain environments requires extensive exploration, but safety requires limiting exploration.
-
25 Apr 2019 2 repositories listedNavigating urban environments represents a complex task for automated vehicles.
-
9 Jun 2025 1 repository listedExisting approaches to language model alignment often treat safety as a tradeoff against helpfulness, which can lead to unacceptable responses in sensitive domains.
-
25 Feb 2025 1 repository listedUtilizing this unified high-level graph and a shared low-level goal-conditioned safe RL policy, we extend this approach to address the multi-agent safe navigation problem.
-
Risk-Averse Reinforcement Learning: An Optimal Transport Perspective on Temporal Difference Learning22 Feb 2025 1 repository listedThe primary goal of reinforcement learning is to develop decision-making policies that prioritize optimal performance, frequently without considering risk or safety.
-
25 Dec 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Offline safe reinforcement learning (OSRL) involves learning a decision-making policy to maximize rewards from a fixed batch of training data to satisfy pre-defined safety constraints.
-
19 Dec 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we propose to learn a policy that generates desirable trajectories and avoids undesirable trajectories.
-
17 Dec 2024 1 repository listedSafe reinforcement learning (RL) is a popular and versatile paradigm to learn reward-maximizing policies with safety guarantees.
-
11 Dec 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)In safe offline reinforcement learning (RL), the objective is to develop a policy that maximizes cumulative rewards while strictly adhering to safety constraints, utilizing only offline data.
-
5 Dec 2024 1 repository listedThe high costs and risks involved in extensive environment interactions hinder the practical application of current online safe reinforcement learning (RL) methods.
-
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning7 Nov 2024 1 repository listedSafe reinforcement learning (RL) is crucial for real-world applications, and multi-agent interactions introduce additional safety challenges.
-
18 Sep 2024 1 repository listedSafety is one of the key issues preventing the deployment of reinforcement learning techniques in real-world robots.
-
14 Aug 2024 1 repository listedThe results show that the simple rules, TreeC, and model predictive control-based methods achieved similar costs, with a difference of only 0.
-
19 Jul 2024 1 repository listed Syntology ran 4 of 8 samples · 4 unverifiedOffline safe reinforcement learning (RL) aims to train a policy that satisfies constraints using a pre-collected dataset.
-
12 Jul 2024 1 repository listedWe analyze Safe BO under the lens of a generalization of active learning with concrete prediction targets where sampling is restricted to an accessible region of the domain, while prediction targets may lie outside this…
-
5 Jun 2024 1 repository listedThe growing complexity of power system management has led to an increased interest in reinforcement learning (RL).
-
31 May 2024 1 repository listedSafe reinforcement learning (RL) is crucial for deploying RL agents in real-world applications, as it aims to maximize long-term rewards while satisfying safety constraints.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections