Papers › Procedural Generalization by Planning with Self-Supervised World Models

Procedural Generalization by Planning with Self-Supervised World Models

2 Nov 2021ICLR 2022 4arXiv:2111.01587archive 2025-07-28

Ankesh Anand, Jacob Walker, Yazhe Li, Eszter Vértes, Julian Schrittwieser, Sherjil Ozair, Théophane Weber, Jessica B. Hamrick

One of the key promises of model-based reinforcement learning is the ability to generalize using an internal model of the world to make predictions in novel environments and tasks. However, the generalization ability of model-based agents is not well understood because existing work has focused on model-free agents when benchmarking generalization. Here, we explicitly measure the generalization ability of model-based agents in comparison to their model-free counterparts. We focus our analysis on MuZero (Schrittwieser et al., 2020), a powerful model-based agent, and evaluate its performance on both procedural and task generalization. We identify three factors of procedural generalization -- planning, self-supervised representation learning, and procedural data diversity -- and show that by combining these techniques, we achieve state-of-the art generalization performance and data efficiency on Procgen (Cobbe et al., 2019). However, we find that these factors do not always provide the same benefits for the task generalization benchmarks in Meta-World (Yu et al., 2019), indicating that transfer remains a challenge and may require different approaches than procedural generalization. Overall, we suggest that building generalizable agents requires moving beyond the single-task, model-free paradigm and towards self-supervised model-based agents that are trained in rich, procedural, multi-task environments.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BenchmarkingMeta-LearningModel-based Reinforcement LearningRepresentation Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Meta-Learning ML10 MZ+Recon Meta-test success rate (zero-shot) 25 #5 of 6 Archive leaderboard report
Meta-Learning ML10 MZ+Recon Meta-train success rate 97.8% #5 of 6 Archive leaderboard report
Meta-Learning ML10 MZ Meta-test success rate (zero-shot) 26.5 #6 of 6 Archive leaderboard report
Meta-Learning ML10 MZ Meta-train success rate 97.6% #6 of 6 Archive leaderboard report
Meta-Learning ML45 MZ+Recon Meta-test success rate (zero-shot) 18.5 #1 of 2 Archive leaderboard report
Meta-Learning ML45 MZ+Recon Meta-train success rate 74.9 #1 of 2 Archive leaderboard report
Meta-Learning ML45 MZ Meta-test success rate (zero-shot) 17.7 #2 of 2 Archive leaderboard report
Meta-Learning ML45 MZ Meta-train success rate 77.2 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Average PoolingBatch NormalizationConvolutionMonte-Carlo Tree SearchMuZeroPrioritized Experience ReplayReLUResidual BlockResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections