Methods › Reinforcement Learning › Offline Reinforcement Learning Methods › URL
Umbrella Reinforcement Learning
URL
Introduced by Egor E. Nuzhin et al. in Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A computationally efficient approach for solving hard nonlinear problems of reinforcement learning (RL). It combines umbrella sampling, from computational physics/chemistry, with optimal control methods. The approach is realized on the basis of neural networks, with the use of policy gradient. It outperforms, by computational efficiency and implementation universality, the available state-of-the-art algorithms, in application to hard RL problems with sparse reward, state traps and lack of terminal states. The proposed approach uses an ensemble of simultaneously acting agents, with a modified reward which includes the ensemble entropy, yielding an optimal exploration-exploitation balance.
Papers archive 2025-07-28
30 shown of 34, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
PhishKey: A Novel Centroid-Based Approach for Enhanced Phishing Detection Using Adaptive HTML Component Extraction 26 Jun 2025 · 0 repositories · arXiv:2506.21106
-
WebGuard++:Interpretable Malicious URL Detection via Bidirectional Fusion of HTML Subgraphs and Multi-Scale Convolutional BERT 24 Jun 2025 · 0 repositories · arXiv:2506.19356
-
Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning 12 Jun 2025 · 0 repositories · arXiv:2506.10629
-
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents 30 May 2025 · 1 repository · arXiv:2505.24878
-
MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection 26 May 2025 · 0 repositories · arXiv:2505.23803
-
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training 22 May 2025 · 0 repositories · arXiv:2505.16570
-
Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection 21 May 2025 · 1 repository · arXiv:2505.16029
-
Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM 21 May 2025 · 1 repository · arXiv:2505.15816Syntology ran 5 of 9 samples · 4 unverified · 4 pointer-only (licence)
-
Do Not Change Me: On Transferring Entities Without Modification in Neural Machine Translation -- a Multilingual Perspective 9 May 2025 · 2 repositories · arXiv:2505.06010
-
Detecting Quishing Attacks with Machine Learning Techniques Through QR Code Analysis 6 May 2025 · 1 repository · arXiv:2505.03451
-
Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning 6 May 2025 · 0 repositories · arXiv:2505.03703
-
Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System 2 May 2025 · 0 repositories · arXiv:2505.01315
-
Phishing URL Detection using Bi-LSTM 29 Apr 2025 · 0 repositories · arXiv:2504.21049
-
A Gradient-Optimized TSK Fuzzy Framework for Explainable Phishing Detection 25 Apr 2025 · 0 repositories · arXiv:2504.18636
-
From Past to Present: A Survey of Malicious URL Detection Techniques, Datasets and Code Repositories 23 Apr 2025 · 0 repositories · arXiv:2504.16449
-
Emerging Cyber Attack Risks of Medical AI Agents 2 Apr 2025 · 0 repositories · arXiv:2504.03759
-
MoonCast: High-Quality Zero-Shot Podcast Generation 18 Mar 2025 · 1 repository · arXiv:2503.14345
-
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search 13 Mar 2025 · 1 repository · arXiv:2503.10582Syntology ran 0 of 13 samples · 13 unverified
-
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms 5 Mar 2025 · 1 repository · arXiv:2503.03644
-
PhishVQC: Optimizing Phishing URL Detection with Correlation Based Feature Selection and Variational Quantum Classifier 3 Mar 2025 · 0 repositories · arXiv:2503.01799
-
Large Engagement Networks for Classifying Coordinated Campaigns and Organic Twitter Trends 1 Mar 2025 · 1 repository · arXiv:2503.00599
-
Tag-Pag: A Dedicated Tool for Systematic Web Page Annotations 22 Feb 2025 · 1 repository · arXiv:2502.16150
-
CyberSentinel: An Emergent Threat Detection System for AI Security 20 Feb 2025 · 0 repositories · arXiv:2502.14966
-
LLM Agents Making Agent Tools 17 Feb 2025 · 1 repository · arXiv:2502.11705
-
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search 7 Feb 2025 · 0 repositories · arXiv:2502.04951
-
A New Dataset and Methodology for Malicious URL Classification 31 Dec 2024 · 1 repository · arXiv:2501.00356
-
GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks 11 Dec 2024 · 1 repository · arXiv:2412.08296
-
Domain-specific Question Answering with Hybrid Search 4 Dec 2024 · 0 repositories · arXiv:2412.03736
-
Stain-aware Domain Alignment for Imbalance Blood Cell Classification 4 Dec 2024 · 1 repository · arXiv:2412.02976
-
Large Multimodal Agents for Accurate Phishing Detection with Enhanced Token Optimization and Cost Reduction 3 Dec 2024 · 0 repositories · arXiv:2412.02301
Tasks archive 2025-07-28
20 shown of 50 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections