{"url":"/task/distributional-reinforcement-learning","name":"Distributional Reinforcement Learning","slug":"distributional-reinforcement-learning","description_markdown":"Value distribution is the distribution of the random return received by a reinforcement learning agent.  it been used for a specific purpose such as implementing risk-aware behaviour. \r\n\r\nWe have random return Z whose expectation is the value Q. This random return is also described by a recursive equation, but one of a distributional nature","categories":[{"name":"Methodology","url":"/area/methodology"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"derived"},"counts":{"papers_tagged":137,"papers_with_code":43,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":0,"subtasks":0,"parent_tasks":0},"benchmarks":[],"datasets":[],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":43,"tagged_in_all":137,"items":[{"url":"/paper/implicit-quantile-networks-for-distributional","title":"Implicit Quantile Networks for Distributional Reinforcement Learning","date":"2018-06-14","arxiv_id":"1806.06923","repositories_listed":19,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/distributional-reinforcement-learning-with-1","title":"Distributional Reinforcement Learning with Quantile Regression","date":"2017-10-27","arxiv_id":"1710.10044","repositories_listed":17,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/fully-parameterized-quantile-function-for","title":"Fully Parameterized Quantile Function for Distributional Reinforcement Learning","date":"2019-11-05","arxiv_id":"1911.02140","repositories_listed":6,"syntology":{"n":10,"n_ran":1,"n_unverified":9,"n_pointer_only":0}},{"url":"/paper/implicit-distributional-reinforcement","title":"Implicit Distributional Reinforcement Learning","date":"2020-07-13","arxiv_id":"2007.06159","repositories_listed":3,"syntology":{"n":6,"n_ran":4,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/quota-the-quantile-option-architecture-for","title":"QUOTA: The Quantile Option Architecture for Reinforcement Learning","date":"2018-11-05","arxiv_id":"1811.02073","repositories_listed":3,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/distributional-reinforcement-learning-for-4","title":"Distributional Reinforcement Learning for Multi-Dimensional Reward Functions","date":"2021-10-26","arxiv_id":"2110.13578","repositories_listed":2,"syntology":null},{"url":"/paper/estimating-risk-and-uncertainty-in-deep","title":"Estimating Risk and Uncertainty in Deep Reinforcement Learning","date":"2019-05-23","arxiv_id":"1905.09638","repositories_listed":2,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/addq-adaptive-distributional-double-q","title":"ADDQ: Adaptive Distributional Double Q-Learning","date":"2025-06-24","arxiv_id":"2506.19478","repositories_listed":1,"syntology":null},{"url":"/paper/rize-regularized-imitation-learning-via","title":"RIZE: Regularized Imitation Learning via Distributional Reinforcement Learning","date":"2025-02-27","arxiv_id":"2502.20089","repositories_listed":1,"syntology":null},{"url":"/paper/tackling-uncertainties-in-multi-agent","title":"Tackling Uncertainties in Multi-Agent Reinforcement Learning through Integration of Agent Termination Dynamics","date":"2025-01-21","arxiv_id":"2501.12061","repositories_listed":1,"syntology":null},{"url":"/paper/beyond-cvar-leveraging-static-spectral-risk","title":"Beyond CVaR: Leveraging Static Spectral Risk Measures for Enhanced Decision-Making in Distributional Reinforcement Learning","date":"2025-01-03","arxiv_id":"2501.02087","repositories_listed":1,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":4}},{"url":"/paper/ex-drl-hedging-against-heavy-losses-with","title":"EX-DRL: Hedging Against Heavy Losses with EXtreme Distributional Reinforcement Learning","date":"2024-08-22","arxiv_id":"2408.12446","repositories_listed":1,"syntology":null},{"url":"/paper/ctd4-a-deep-continuous-distributional-actor","title":"CTD4 -- A Deep Continuous Distributional Actor-Critic Agent with a Kalman Fusion of Multiple Critics","date":"2024-05-04","arxiv_id":"2405.02576","repositories_listed":1,"syntology":null},{"url":"/paper/a-distributional-analogue-to-the-successor","title":"A Distributional Analogue to the Successor Representation","date":"2024-02-13","arxiv_id":"2402.08530","repositories_listed":1,"syntology":{"n":4,"n_ran":2,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/echoes-of-socratic-doubt-embracing","title":"Echoes of Socratic Doubt: Embracing Uncertainty in Calibrated Evidential Reinforcement Learning","date":"2024-02-11","arxiv_id":"2402.07107","repositories_listed":1,"syntology":null},{"url":"/paper/distributional-off-policy-evaluation-with","title":"Distributional Off-policy Evaluation with Bellman Residual Minimization","date":"2024-02-02","arxiv_id":"2402.01900","repositories_listed":1,"syntology":null},{"url":"/paper/a-robust-quantile-huber-loss-with","title":"A Robust Quantile Huber Loss With Interpretable Parameter Adjustment In Distributional Reinforcement Learning","date":"2024-01-04","arxiv_id":"2401.02325","repositories_listed":1,"syntology":null},{"url":"/paper/distributional-bellman-operators-over-mean","title":"Distributional Bellman Operators over Mean Embeddings","date":"2023-12-09","arxiv_id":"2312.07358","repositories_listed":1,"syntology":null},{"url":"/paper/estimation-and-inference-in-distributional","title":"Estimation and Inference in Distributional Reinforcement Learning","date":"2023-09-29","arxiv_id":"2309.17262","repositories_listed":1,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":4}},{"url":"/paper/value-distributional-model-based","title":"Value-Distributional Model-Based Reinforcement Learning","date":"2023-08-12","arxiv_id":"2308.06590","repositories_listed":1,"syntology":null},{"url":"/paper/variance-control-for-distributional","title":"Variance Control for Distributional Reinforcement Learning","date":"2023-07-30","arxiv_id":"2307.16152","repositories_listed":1,"syntology":null},{"url":"/paper/distributional-model-equivalence-for-risk","title":"Distributional Model Equivalence for Risk-Sensitive Reinforcement Learning","date":"2023-07-04","arxiv_id":"2307.01708","repositories_listed":1,"syntology":null},{"url":"/paper/the-benefits-of-being-distributional-small-1","title":"The Benefits of Being Distributional: Small-Loss Bounds for Reinforcement Learning","date":"2023-05-25","arxiv_id":"2305.15703","repositories_listed":1,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":1}},{"url":"/paper/trustworthy-reinforcement-learning-for","title":"Constrained Reinforcement Learning using Distributional Representation for Trustworthy Quadrotor UAV Tracking Control","date":"2023-02-22","arxiv_id":"2302.11694","repositories_listed":1,"syntology":null},{"url":"/paper/distributional-constrained-reinforcement","title":"Distributional constrained reinforcement learning for supply chain optimization","date":"2023-02-03","arxiv_id":"2302.01727","repositories_listed":1,"syntology":null},{"url":"/paper/efficient-trust-region-based-safe","title":"Trust Region-Based Safe Distributional Reinforcement Learning for Multiple Constraints","date":"2023-01-26","arxiv_id":"2301.10923","repositories_listed":1,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/risk-sensitive-policy-with-distributional","title":"Risk-Sensitive Policy with Distributional Reinforcement Learning","date":"2022-12-30","arxiv_id":"2212.14743","repositories_listed":1,"syntology":null},{"url":"/paper/intelligent-resource-allocation-in-joint","title":"Intelligent Resource Allocation in Joint Radar-Communication With Graph Neural Networks","date":"2022-10-17","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/ign-implicit-generative-networks","title":"IGN : Implicit Generative Networks","date":"2022-06-13","arxiv_id":"2206.05860","repositories_listed":1,"syntology":null},{"url":"/paper/gamma-and-vega-hedging-using-deep","title":"Gamma and Vega Hedging Using Deep Distributional Reinforcement Learning","date":"2022-05-10","arxiv_id":"2205.05614","repositories_listed":1,"syntology":null}],"syntology_records":11,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}