Browse State-of-the-Art › Hierarchical Reinforcement Learning
Hierarchical Reinforcement Learning
111 papers with code · 1 benchmark · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Ant + Maze (1 row) | STAR | Reconciling Spatial and Temporal Abstractions for Goal Representation | code | Syntology ran 5 of 11 samples · 6 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 111 papers with code (384 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Dec 2010 15 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedWe present a tutorial on Bayesian optimization, a method of finding the maximum of expensive cost functions.
-
21 May 2018 12 repositories listedIn this paper, we study how we can develop HRL algorithms that are general, in that they do not make onerous additional assumptions beyond standard RL algorithms, and efficient, in the sense that they can be used with…
-
2 Oct 2018 6 repositories listedWe study the problem of representation learning in goal-conditioned hierarchical reinforcement learning.
-
21 May 1999 5 repositories listedThe paper presents an online model-free learning algorithm, MAXQ-Q, and proves that it converges wih probability 1 to a kind of locally-optimal policy known as a recursively optimal policy, even in the presence of the…
-
29 Apr 2020 4 repositories listedExisting approaches usually employ a flat policy structure that treat all symptoms and diseases equally for action making.
-
4 Dec 2017 4 repositories listedHierarchical agents have the potential to solve sequential decision making tasks with greater sample efficiency than their non-hierarchical counterparts because hierarchical agents can break down tasks into sets of…
-
10 Feb 2025 3 repositories listedWe train our ReasonFlux-32B model with only 8 GPUs and introduces three innovations: (i) a structured and generic thought template library, containing around 500 high-level thought templates capable of generalizing to…
-
19 Jul 2022 3 repositories listedDue to this one-to-many dilemma, enlarged action space and ignoring logical relationship between entity and relation increase the difficulty of learning.
-
12 Oct 2024 2 repositories listedGoal-conditioned hierarchical reinforcement learning (HRL) decomposes complex reaching tasks into a sequence of simple subgoal-conditioned tasks, showing significant promise for addressing long-horizon planning in…
-
25 Jun 2024 2 repositories listed Syntology ran 3 of 7 samples · 4 unverified · 7 pointer-only (licence)Specifically, the whole VNE process is decomposed into an upper-level policy for deciding whether to admit the arriving VNR or not and a lower-level policy for allocating resources of the physical network to meet the…
-
26 Aug 2022 2 repositories listedMotivated by this, we investigate developing a learning-based over-sampling algorithm to optimize the classification performance, which is a challenging task because of the huge and hierarchical decision space.
-
22 Mar 2021 2 repositories listedThis problem is referred to as hierarchical imitation learning and can be handled as an inference problem in a Hidden Markov Model, which is done via an Expectation-Maximization type algorithm.
-
11 Jun 2020 2 repositories listedWe test HDNO on MultiWoz 2.
-
12 Nov 2019 2 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedFurthermore, to approximate solutions to constrained combinatorial optimization problems such as the TSP with time windows, we train hierarchical GPNs (HGPNs) using RL, which learns a hierarchical policy to find an…
-
31 Oct 2019 2 repositories listedFairness is essential for human society, contributing to stability and productivity.
-
22 Nov 2018 2 repositories listedIn hierarchical reinforcement learning a major challenge is determining appropriate low-level policies.
-
9 Nov 2018 2 repositories listedThe whole extraction process is decomposed into a hierarchy of two-level RL policies for relation detection and entity extraction respectively, so that it is more feasible and natural to deal with overlapping relations.
-
10 Apr 2017 2 repositories listedThen a high-level policy is trained on top of these skills, providing a significant improvement of the exploration and allowing to tackle sparse rewards in the downstream tasks.
-
11 Jun 2025 1 repository listedTo that end, we introduce SkillBlender, a novel hierarchical reinforcement learning framework for versatile humanoid loco-manipulation.
-
9 Jun 2025 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedTo overcome these challenges, we recast multi-LLM coordination as an incomplete-information game and seek a Bayesian Nash equilibrium (BNE), in which each agent optimally responds to its probabilistic beliefs about the…
-
26 May 2025 1 repository listedWhile showing sophisticated reasoning abilities, large language models (LLMs) still struggle with long-horizon decision-making tasks due to deficient exploration and long-term credit assignment, especially in…
-
7 May 2025 1 repository listed Syntology ran 0 of 7 samples · 7 unverifiedIn collaborative tasks, autonomous agents fall short of humans in their capability to quickly adapt to new and unfamiliar teammates.
-
6 Apr 2025 1 repository listedWe introduce a novel hierarchical reinforcement learning (HRL) framework that performs top-down recursive planning via learned subgoals, successfully applied to the complex combinatorial puzzle game Sokoban.
-
19 Mar 2025 1 repository listedOur empirical evaluation on a suite of procedurally generated continuous control environments demonstrates that our approach outperforms existing hierarchical reinforcement learning methods in terms of sample…
-
27 Feb 2025 1 repository listedWe provide analysis of the design choices of the reward models and policy, and show the efficacy of μCode at utilizing the execution feedback.
-
21 Feb 2025 1 repository listedHierarchical organization is fundamental to biological systems and human societies, yet artificial intelligence systems often rely on monolithic architectures that limit adaptability and scalability.
-
20 Dec 2024 1 repository listedOur approach continually learns and maintains an interpretable state abstraction, and uses it to invent high-level options with abstract symbolic representations.
-
10 Oct 2024 1 repository listedTo address this, we integrate meta-learning into HRL to enhance the agent's ability to learn and adapt hierarchical policies swiftly.
-
20 Sep 2024 1 repository listedIn recent years, robots and autonomous systems have become increasingly integral to our daily lives, offering solutions to complex problems across various domains.
-
24 Aug 2024 1 repository listedLearning path recommendation aims to provide learners with a reasonable order of items to achieve their learning goals.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections