Browse State-of-the-Art › Continuous Control

Continuous Control

494 papers with code · 73 benchmarks · 10 datasets archive 2025-07-28

Computer VisionPlaying GamesRobots

Continuous control in the context of playing games, especially within artificial intelligence (AI) and machine learning (ML), refers to the ability to make a series of smooth, ongoing adjustments or actions to control a game or a simulation. This is in contrast to discrete control, where the actions are limited to a set of specific, distinct choices. Continuous control is crucial in environments where precision, timing, and the magnitude of actions matter, such as driving a car in a racing game, controlling a character in a simulation, or managing the flight of an aircraft in a flight simulator.

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

73 leaderboard tables shown for this task, 73 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 73 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
PyBullet Ant (8 rows) SAC gSDE Smooth Exploration for Robotic Reinforcement Learning code Syntology ran 0 of 1 samples · 1 unverified Compare
PyBullet HalfCheetah (8 rows) SAC Smooth Exploration for Robotic Reinforcement Learning code Syntology ran 0 of 1 samples · 1 unverified Compare
PyBullet Hopper (8 rows) SAC gSDE Smooth Exploration for Robotic Reinforcement Learning code Syntology ran 0 of 1 samples · 1 unverified Compare
PyBullet Walker2D (8 rows) SAC gSDE Smooth Exploration for Robotic Reinforcement Learning code Syntology ran 0 of 1 samples · 1 unverified Compare
Lunar Lander (OpenAI Gym) (5 rows) SAC Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement... code Syntology ran 91 of 148 samples · 57 unverified Compare
DeepMind Cheetah Run (Images) (4 rows) DreamerV1 — — — Compare
cartpole.balance_sparse (2 rows) SMuZero Learning and Planning in Complex Action Spaces code — Compare
cartpole.swingup (2 rows) SMuZero Learning and Planning in Complex Action Spaces code — Compare
cheetah.run (2 rows) SMuZero Learning and Planning in Complex Action Spaces code — Compare
DeepMind Cup Catch (Images) (2 rows) DrQ Image Augmentation Is All You Need: Regularizing Deep... code Syntology ran 6 of 10 samples · 4 unverified Compare
DeepMind Walker Walk (Images) (2 rows) DrQ Image Augmentation Is All You Need: Regularizing Deep... code Syntology ran 6 of 10 samples · 4 unverified Compare
finger.turn_hard (2 rows) SMuZero Learning and Planning in Complex Action Spaces code — Compare
walker.stand (2 rows) SMuZero Learning and Planning in Complex Action Spaces code — Compare
walker.walk (2 rows) SMuZero Learning and Planning in Complex Action Spaces code — Compare
2D Walker (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Acrobot (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Acrobot (limited sensors) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Acrobot (noisy observations) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
acrobot.swingup (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Acrobot (system identifications) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Ant (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Ant + Gathering (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Ant + Maze (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Ball in cup, catch (DMControl500k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Ball in cup, catch (DMControl100k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
ball_in_cup.catch (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Cart-Pole Balancing (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Cart-Pole Balancing (limited sensors) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Cart-Pole Balancing (noisy observations) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Cart-Pole Balancing (system identifications) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Cart Pole (OpenAI Gym) (1 row) MAC Mean Actor Critic code — Compare
cartpole.balance (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Cartpole, swingup (DMControl500k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Cartpole, swingup (DMControl100k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
cartpole.swingup_sparse (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Cheetah, run (DMControl500k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Cheetah, run (DMControl100k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Double Inverted Pendulum (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Finger, spin (DMControl500k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Finger, spin (DMControl100k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
finger.spin (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
finger.turn_easy (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
fish.swim (1 row) MuZero Unplugged Online and Offline Reinforcement Learning by Planning with a Learned Model code — Compare
Full Humanoid (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Half-Cheetah (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Hopper (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
hopper.hop (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
hopper.stand (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
humanoid.run (1 row) MuZero Unplugged Online and Offline Reinforcement Learning by Planning with a Learned Model code — Compare
Inverted Pendulum (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Inverted Pendulum (limited sensors) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Inverted Pendulum (system identifications) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Inverted Pendulum (noisy observations) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
manipulator.insert_ball (1 row) MuZero Unplugged Online and Offline Reinforcement Learning by Planning with a Learned Model code — Compare
manipulator.insert_peg (1 row) MuZero Unplugged Online and Offline Reinforcement Learning by Planning with a Learned Model code — Compare
Mountain Car (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Mountain Car (limited sensors) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Mountain Car (noisy observations) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Mountain Car (system identifications) (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
pendulum.swingup (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
quadruped.run (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
quadruped.walk (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Reacher, easy (DMControl500k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Reacher, easy (DMControl100k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
reacher.easy (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
reacher.hard (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Simple Humanoid (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Swimmer (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Swimmer + Gathering (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
Swimmer + Maze (1 row) TRPO Benchmarking Deep Reinforcement Learning for Continuous Control code Syntology ran 0 of 4 samples · 4 unverified Compare
walker.run (1 row) SMuZero Learning and Planning in Complex Action Spaces code — Compare
Walker, walk (DMControl500k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
Walker, walk (DMControl100k) (1 row) CURL CURL: Contrastive Unsupervised Representations for Reinforcement Learning code Syntology ran 6 of 8 samples · 2 unverified Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

10 datasets whose archive record lists this task, ordered by the archive's paper count.

Subtasks archive 2025-07-28

3 subtasks in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 494 papers with code (1,161 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 26 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections