Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › URL

Umbrella Reinforcement Learning

URL

34 papers tagged archive 2025-07-28

Introduced by Egor E. Nuzhin et al. in Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

A computationally efficient approach for solving hard nonlinear problems of reinforcement learning (RL). It combines umbrella sampling, from computational physics/chemistry, with optimal control methods. The approach is realized on the basis of neural networks, with the use of policy gradient. It outperforms, by computational efficiency and implementation universality, the available state-of-the-art algorithms, in application to hard RL problems with sparse reward, state traps and lack of terminal states. The proposed approach uses an ensemble of simultaneously acting agents, with a modified reward which includes the ensemble entropy, yielding an optimal exploration-exploitation balance.

PaperSource

Papers archive 2025-07-28

30 shown of 34, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 50 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Contrastive Learning3
Benchmarking2
Classification2
Zero-Shot Learning2
zero-shot-classification2
3D Pedestrian Tracking1
Anomaly Detection1
Binary Classification1
Blocking1
Code Generation1
Computational Efficiency1
Decoder1
Disentanglement1
Diversity1
Efficient Exploration1
Feature Importance1
Graph Classification1
Graph Neural Network1
Image Retrieval1
Image Segmentation1

Usage over time archive 2025-07-28

Papers per year tagged with URL: 2024 to 2025, peak 25 25 0 2024: 9 papers 2024 2025: 25 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (34 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Offline Reinforcement Learning Methods

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections