Browse State-of-the-Art › OpenAI Gym
OpenAI Gym
179 papers with code · 17 benchmarks · 3 datasets archive 2025-07-28
An open-source toolkit from OpenAI that implements several Reinforcement Learning benchmarks including: classic control, Atari, Robotics and MuJoCo tasks.
(Description by Evolutionary learning of interpretable decision trees)
(Image Credit: OpenAI Gym)
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
17 leaderboard tables shown for this task, 17 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 17 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 179 papers with code (382 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Jul 2017 188 repositories listed Syntology ran 99 of 176 samples · 77 unverified · 94 pointer-only (licence)We propose a new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using…
-
9 Sep 2015 161 repositories listed Syntology ran 158 of 306 samples · 148 unverified · 163 pointer-only (licence)We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain.
-
4 Jan 2018 86 repositories listed Syntology ran 91 of 148 samples · 57 unverified · 66 pointer-only (licence)A platform for Applied Reinforcement Learning (Applied RL)
-
26 Feb 2018 67 repositories listed Syntology ran 9 of 36 samples · 27 unverified · 20 pointer-only (licence)In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies.
-
26 Feb 2018 28 repositories listedThe purpose of this technical report is two-fold.
-
2 Jun 2021 20 repositories listed Syntology ran 17 of 26 samples · 9 unverified · 6 pointer-only (licence)In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling.
-
1 Oct 2019 6 repositories listedIn this paper, we aim to develop a simple and scalable reinforcement learning algorithm that uses standard supervised learning methods as subroutines.
-
23 Jul 2015 5 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedDeep Reinforcement Learning has yielded proficient controllers for complex tasks.
-
23 Jun 2023 4 repositories listedDespite its simplicity this baseline is competitive with meta-learning methods on a variety of conditions and is able to imitate target policies trained on unseen variations of the original environment.
-
5 May 2018 4 repositories listedDeep reinforcement learning has shown its success in game playing.
-
3 Mar 2021 3 repositories listedRobotic simulators are crucial for academic research and education as well as the development of safety-critical applications.
-
13 Jul 2020 3 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedTo improve the sample efficiency of policy-gradient based reinforcement learning algorithms, we propose implicit distributional actor-critic (IDAC) that consists of a distributional critic, built on two deep generator…
-
8 Oct 2019 3 repositories listed Syntology ran 4 of 9 samples · 5 unverified · 3 pointer-only (licence)TorchBeast is a platform for reinforcement learning (RL) research in PyTorch.
-
21 May 2019 3 repositories listedThis objective encourages the agent to maximize the expected return, as well as to achieve more diverse goals.
-
13 May 2024 2 repositories listedIn this work, we introduce two novel methods, Decision Mamba (DM) and Hierarchical Decision Mamba (HDM), aimed at enhancing the performance of the Transformer models.
-
4 Jun 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)In the field of reinforcement learning (RL), representation learning is a proven tool for complex image-based tasks, but is often overlooked for environments with low-level states, such as physical control problems.
-
30 Nov 2022 2 repositories listedWe introduce MO-Gym, an extensible library containing a diverse set of multi-objective reinforcement learning environments.
-
11 Nov 2022 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe present pyRDDLGym, a Python framework for auto-generation of OpenAI Gym environments from RDDL declerative description.
-
15 Sep 2022 2 repositories listedThis paper presents COOL-MC, a tool that integrates state-of-the-art reinforcement learning (RL) and model checking.
-
11 Aug 2022 2 repositories listedAdopting reasonable strategies is challenging but crucial for an intelligent agent with limited resources working in hazardous, unstructured, and dynamic environments to improve the system's utility, decrease the…
-
24 Feb 2022 2 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedWe utilize hybrid quantum deep reinforcement learning to learn navigation tasks for a simple, wheeled robot in simulated environments of increasing complexity.
-
21 Jul 2021 2 repositories listedThis paper is an initial endeavor to bridge the gap between powerful Deep Reinforcement Learning methodologies and the problem of exploration/coverage of unknown terrains.
-
12 May 2021 2 repositories listedThis work re-implements the OpenAI Gym multi-goal robotic manipulation environment, originally based on the commercial Mujoco engine, onto the open-source Pybullet engine.
-
9 Apr 2021 2 repositories listedNitrogen fertilizers have a detrimental effect on the environment, which can be reduced by optimizing fertilizer management strategies.
-
10 Feb 2021 2 repositories listedUsing a model of the environment, reinforcement learning agents can plan their future moves and achieve superhuman performance in board games like Chess, Shogi, and Go, while remaining relatively sample-efficient.
-
8 Feb 2021 2 repositories listedAutomatic programming, the task of generating computer programs compliant with a specification without a human developer, is usually tackled either via genetic programming methods based on mutation and recombination of…
-
29 Dec 2020 2 repositories listedThis paper is a study of reinforcement learning (RL) as an optimal-control strategy for control of nonlinear valves.
-
14 Nov 2020 2 repositories listedFurther, we evaluate a variety of algorithms on these tasks and highlight challenges for reinforcement learning algorithms, including dealing with a state representation that has a high intrinsic dimensionality and is…
-
11 Nov 2020 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe present Ecole, a new library to simplify machine learning research for combinatorial optimization.
-
9 Oct 2020 2 repositories listedEpidemiologists model the dynamics of epidemics in order to propose control strategies based on pharmaceutical and non-pharmaceutical interventions (contact limitation, lock down, vaccination, etc).
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections