Papers › Learning and Planning in Complex Action Spaces

Learning and Planning in Complex Action Spaces

13 Apr 2021arXiv:2104.06303archive 2025-07-28

Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain, Simon Schmitt, David Silver

Many important real-world problems have action spaces that are high-dimensional, continuous or both, making full enumeration of all possible actions infeasible. Instead, only small subsets of actions can be sampled for the purpose of policy evaluation and improvement. In this paper, we propose a general framework to reason in a principled way about policy evaluation and improvement over such sampled action subsets. This sample-based policy iteration framework can in principle be applied to any reinforcement learning algorithm based upon policy iteration. Concretely, we propose Sampled MuZero, an extension of the MuZero algorithm that is able to learn in domains with arbitrarily complex action spaces by planning over sampled actions. We demonstrate this approach on the classical board game of Go and on two continuous control benchmark domains: DeepMind Control Suite and Real-World RL Suite.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

opendilab/LightZero mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Continuous ControlGame of Gocontinuous-control

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Continuous Control acrobot.swingup SMuZero Return 417.52 #1 of 1 Archive leaderboard report
Continuous Control ball_in_cup.catch SMuZero Return 977.38 #1 of 1 Archive leaderboard report
Continuous Control cartpole.balance SMuZero Return 984.86 #1 of 1 Archive leaderboard report
Continuous Control cartpole.balance_sparse SMuZero Return 998.14 #1 of 2 Archive leaderboard report
Continuous Control cartpole.swingup SMuZero Return 868.87 #1 of 2 Archive leaderboard report
Continuous Control cartpole.swingup_sparse SMuZero Return 846.91 #1 of 1 Archive leaderboard report
Continuous Control cheetah.run SMuZero Return 914.39 #1 of 2 Archive leaderboard report
Continuous Control finger.spin SMuZero Return 986.38 #1 of 1 Archive leaderboard report
Continuous Control finger.turn_easy SMuZero Return 972.53 #1 of 1 Archive leaderboard report
Continuous Control finger.turn_hard SMuZero Return 963.07 #1 of 2 Archive leaderboard report
Continuous Control hopper.hop SMuZero Return 528.24 #1 of 1 Archive leaderboard report
Continuous Control hopper.stand SMuZero Return 926.5 #1 of 1 Archive leaderboard report
Continuous Control pendulum.swingup SMuZero Return 837.76 #1 of 1 Archive leaderboard report
Continuous Control quadruped.run SMuZero Return 923.54 #1 of 1 Archive leaderboard report
Continuous Control quadruped.walk SMuZero Return 933.77 #1 of 1 Archive leaderboard report
Continuous Control reacher.easy SMuZero Return 982.26 #1 of 1 Archive leaderboard report
Continuous Control reacher.hard SMuZero Return 971.53 #1 of 1 Archive leaderboard report
Continuous Control walker.run SMuZero Return 931.06 #1 of 1 Archive leaderboard report
Continuous Control walker.stand SMuZero Return 987.79 #1 of 2 Archive leaderboard report
Continuous Control walker.walk SMuZero Return 975.46 #1 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Average PoolingBatch NormalizationConvolutionMonte-Carlo Tree SearchMuZeroPrioritized Experience ReplayReLUResidual BlockResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections