Browse State-of-the-Art › Continuous Control
Continuous Control
494 papers with code · 73 benchmarks · 10 datasets archive 2025-07-28
Continuous control in the context of playing games, especially within artificial intelligence (AI) and machine learning (ML), refers to the ability to make a series of smooth, ongoing adjustments or actions to control a game or a simulation. This is in contrast to discrete control, where the actions are limited to a set of specific, distinct choices. Continuous control is crucial in environments where precision, timing, and the magnitude of actions matter, such as driving a car in a racing game, controlling a character in a simulation, or managing the flight of an aircraft in a flight simulator.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
73 leaderboard tables shown for this task, 73 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 73 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 494 papers with code (1,161 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Jul 2017 188 repositories listed Syntology ran 99 of 176 samples · 77 unverified · 94 pointer-only (licence)We propose a new family of policy gradient methods for reinforcement learning, which alternate between sampling data through interaction with the environment, and optimizing a "surrogate" objective function using…
-
9 Sep 2015 161 repositories listed Syntology ran 158 of 306 samples · 148 unverified · 163 pointer-only (licence)We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain.
-
9 Sep 2015 161 repositories listed Syntology ran 158 of 306 samples · 148 unverified · 163 pointer-only (licence)We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain.
-
4 Jan 2018 86 repositories listed Syntology ran 91 of 148 samples · 57 unverified · 66 pointer-only (licence)A platform for Applied Reinforcement Learning (Applied RL)
-
26 Feb 2018 67 repositories listed Syntology ran 9 of 36 samples · 27 unverified · 20 pointer-only (licence)In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies.
-
26 Feb 2018 28 repositories listedThe purpose of this technical report is two-fold.
-
26 Feb 2018 28 repositories listedThe purpose of this technical report is two-fold.
-
19 Mar 2018 26 repositories listed Syntology ran 3 of 15 samples · 12 unverified · 3 pointer-only (licence)A common belief in model-free reinforcement learning is that methods based on random search in the parameter space of policies exhibit significantly worse sample complexity than those that explore the space of actions.
-
19 Mar 2018 26 repositories listed Syntology ran 3 of 15 samples · 12 unverified · 3 pointer-only (licence)A common belief in model-free reinforcement learning is that methods based on random search in the parameter space of policies exhibit significantly worse sample complexity than those that explore the space of actions.
-
3 Dec 2019 21 repositories listed Syntology ran 43 of 62 samples · 19 unverified · 9 pointer-only (licence)Learned world models summarize an agent's experience to facilitate learning complex behaviors.
-
8 Jun 2020 18 repositories listed Syntology ran 24 of 34 samples · 10 unverified · 5 pointer-only (licence)We theoretically show that CQL produces a lower bound on the value of the current policy and that it can be incorporated into a policy learning procedure with theoretical improvement guarantees.
-
8 Jun 2020 18 repositories listed Syntology ran 24 of 34 samples · 10 unverified · 5 pointer-only (licence)We theoretically show that CQL produces a lower bound on the value of the current policy and that it can be incorporated into a policy learning procedure with theoretical improvement guarantees.
-
8 Jun 2015 17 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural…
-
8 Jun 2015 17 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural…
-
22 Apr 2016 15 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedRecently, researchers have made significant progress combining the advances in deep learning for learning feature representations with reinforcement learning.
-
22 Apr 2016 15 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedRecently, researchers have made significant progress combining the advances in deep learning for learning feature representations with reinforcement learning.
-
7 Dec 2018 10 repositories listed Syntology ran 13 of 14 samples · 1 unverified · 9 pointer-only (licence)Many practical applications of reinforcement learning constrain agents to learn from a fixed batch of data which has already been gathered, without offering further possibility for data collection.
-
7 Dec 2018 10 repositories listed Syntology ran 13 of 14 samples · 1 unverified · 9 pointer-only (licence)Many practical applications of reinforcement learning constrain agents to learn from a fixed batch of data which has already been gathered, without offering further possibility for data collection.
-
6 Jun 2017 10 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Combining parameter noise with traditional RL methods allows to combine the best of both worlds.
-
6 Jun 2017 10 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Combining parameter noise with traditional RL methods allows to combine the best of both worlds.
-
1 Jul 2019 9 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedDeep reinforcement learning (RL) algorithms can use high-capacity deep networks to learn directly from image observations.
-
1 Jul 2019 9 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedDeep reinforcement learning (RL) algorithms can use high-capacity deep networks to learn directly from image observations.
-
12 Nov 2018 9 repositories listed Syntology ran 2 of 7 samples · 5 unverified · 1 pointer-only (licence)Planning has been very successful for control tasks with known environment dynamics.
-
12 Nov 2018 9 repositories listed Syntology ran 2 of 7 samples · 5 unverified · 1 pointer-only (licence)Planning has been very successful for control tasks with known environment dynamics.
-
20 Jul 2021 8 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedWe present DrQ-v2, a model-free reinforcement learning (RL) algorithm for visual continuous control.
-
20 Jul 2021 8 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedWe present DrQ-v2, a model-free reinforcement learning (RL) algorithm for visual continuous control.
-
2 Jan 2018 8 repositories listedThe DeepMind Control Suite is a set of continuous control tasks with a standardised structure and interpretable rewards, intended to serve as performance benchmarks for reinforcement learning agents.
-
2 Jan 2018 8 repositories listedThe DeepMind Control Suite is a set of continuous control tasks with a standardised structure and interpretable rewards, intended to serve as performance benchmarks for reinforcement learning agents.
-
17 Aug 2017 8 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we propose to apply trust region optimization to deep reinforcement learning using a recently proposed Kronecker-factored approximation to the curvature.
-
17 Aug 2017 8 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we propose to apply trust region optimization to deep reinforcement learning using a recently proposed Kronecker-factored approximation to the curvature.
Syntology lines on 26 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections