Browse State-of-the-Art › Distributional Reinforcement Learning
Distributional Reinforcement Learning
43 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Value distribution is the distribution of the random return received by a reinforcement learning agent. it been used for a specific purpose such as implementing risk-aware behaviour.
We have random return Z whose expectation is the value Q. This random return is also described by a recursive equation, but one of a distributional nature
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 43 papers with code (137 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Jun 2018 19 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn this work, we build on recent advances in distributional reinforcement learning to give a generally applicable, flexible, and state-of-the-art distributional variant of DQN.
-
27 Oct 2017 17 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)In this paper, we build on recent work advocating a distributional approach to reinforcement learning in which the distribution over returns is modeled explicitly instead of only estimating the mean.
-
5 Nov 2019 6 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedThe key challenge in practical distributional RL algorithms lies in how to parameterize estimated distributions so as to better approximate the true continuous distribution.
-
13 Jul 2020 3 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedTo improve the sample efficiency of policy-gradient based reinforcement learning algorithms, we propose implicit distributional actor-critic (IDAC) that consists of a distributional critic, built on two deep generator…
-
5 Nov 2018 3 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn this paper, we propose the Quantile Option Architecture (QUOTA) for exploration based on recent advances in distributional reinforcement learning (RL).
-
26 Oct 2021 2 repositories listedTo fully inherit the benefits of distributional RL and hybrid reward architectures, we introduce Multi-Dimensional Distributional DQN (MD3QN), which extends distributional RL to model the joint return distribution from…
-
23 May 2019 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedReinforcement learning agents are faced with two types of uncertainty.
-
24 Jun 2025 1 repository listedWe propose an easy to implement method built on top of distributional reinforcement learning (DRL) algorithms to deal with the overestimation in a locally adaptive way.
-
27 Feb 2025 1 repository listedWe introduce a novel Inverse Reinforcement Learning (IRL) approach that overcomes limitations of fixed reward assignments and constrained flexibility in implicit reward regularization.
-
21 Jan 2025 1 repository listedMulti-Agent Reinforcement Learning (MARL) has gained significant traction for solving complex real-world tasks, but the inherent stochasticity and uncertainty in these environments pose substantial challenges to…
-
3 Jan 2025 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Additionally, we provide a clear interpretation of the learned policy by leveraging the distribution of returns in DRL and the decomposition of static coherent risk measures.
-
22 Aug 2024 1 repository listedHowever, these risk measures depend on the accurate estimation of extreme quantiles in the loss distribution's tail, which can be imprecise in QR-based DRL due to the rarity and extremity of tail data, as highlighted in…
-
CTD4 -- A Deep Continuous Distributional Actor-Critic Agent with a Kalman Fusion of Multiple Critics4 May 2024 1 repository listedCategorical Distributional Reinforcement Learning (CDRL) has demonstrated superior sample efficiency in learning complex tasks compared to conventional Reinforcement Learning (RL) approaches.
-
13 Feb 2024 1 repository listed Syntology ran 2 of 4 samples · 2 unverifiedThis paper contributes a new approach for distributional reinforcement learning which elucidates a clean separation of transition structure and reward in the learning process.
-
11 Feb 2024 1 repository listedWe present a novel statistical approach to incorporating uncertainty awareness in model-free distributional reinforcement learning involving quantile regression-based deep Q networks.
-
2 Feb 2024 1 repository listedIn contrast, we study the more manageable expectation-extended statistical distances and provide a novel theoretical justification on their validity for learning the return distribution.
-
4 Jan 2024 1 repository listedDistributional Reinforcement Learning (RL) estimates return distribution mainly by learning quantile values via minimizing the quantile Huber loss function, entailing a threshold parameter often selected heuristically…
-
9 Dec 2023 1 repository listedWe propose a novel algorithmic framework for distributional reinforcement learning, based on learning finite-dimensional mean embeddings of return distributions.
-
29 Sep 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)This implies the distributional policy evaluation problem can be solved with sample efficiency.
-
12 Aug 2023 1 repository listedQuantifying uncertainty about a policy's long-term performance is important to solve sequential decision-making tasks.
-
30 Jul 2023 1 repository listedAlthough distributional reinforcement learning (DRL) has been widely examined in the past few years, very few studies investigate the validity of the obtained Q-function estimator in the distributional setting.
-
4 Jul 2023 1 repository listedWe consider the problem of learning models for risk-sensitive reinforcement learning.
-
25 May 2023 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)In online RL, we propose a DistRL algorithm that constructs confidence sets using maximum likelihood estimation.
-
22 Feb 2023 1 repository listedWe propose a novel, interpretable trajectory tracker integrating a Distributional Reinforcement Learning disturbance estimator for unknown aerodynamic effects with a Stochastic Model Predictive Controller (SMPC).
-
3 Feb 2023 1 repository listedWe introduce Distributional Constrained Policy Optimization (DCPO), a novel approach for reliable constraint satisfaction in RL.
-
26 Jan 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedIn safety-critical robotic tasks, potential failures must be reduced, and multiple constraints must be met, such as avoiding collisions, limiting energy consumption, and maintaining balance.
-
30 Dec 2022 1 repository listedClassical reinforcement learning (RL) techniques are generally concerned with the design of decision-making policies driven by the maximisation of the expected outcome.
-
17 Oct 2022 1 repository listedIn this paper, we propose a framework for intelligent vehicles to conduct JRC, with minimal prior knowledge of the system model and a tunable performance balance, in an environment where surrounding vehicles execute…
-
13 Jun 2022 1 repository listedIn this work, we build recent advances in distributional reinforcement learning to give a state-of-art distributional variant of the model based on the IQN.
-
10 May 2022 1 repository listedWe show how D4PG can be used in conjunction with quantile regression to develop a hedging strategy for a trader responsible for derivatives that arrive stochastically and depend on a single underlying asset.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections