{"url":"/method/url","slug":"url","name":"URL","full_name":"Umbrella Reinforcement Learning","full_name_withheld":false,"description_markdown":"A computationally efficient approach for solving hard nonlinear problems of reinforcement learning (RL). It combines umbrella sampling, from computational physics/chemistry, with optimal control methods. The approach is realized on the basis of neural networks, with the use of policy gradient. It outperforms, by computational efficiency and implementation universality, the available state-of-the-art algorithms, in application to hard RL problems with sparse reward, state traps and lack of terminal states. The proposed approach uses an ensemble of simultaneously acting agents, with a modified reward which includes the ensemble entropy, yielding an optimal exploration-exploitation balance.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems","paper":"/paper/umbrella-reinforcement-learning","first_author":"Egor E. Nuzhin","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/umbrella-reinforcement-learning"},"source":{"url":"https://arxiv.org/abs/2411.14117v1","title":"Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Offline Reinforcement Learning Methods","url":"/methods/category/offline-reinforcement-learning-methods","pwc_aliases":[]}],"n_papers_tagged":34,"archive_num_papers":35,"papers_newest_first":[{"paper":null,"title":"PhishKey: A Novel Centroid-Based Approach for Enhanced Phishing Detection Using Adaptive HTML Component Extraction","date":"2025-06-26","arxiv_id":"2506.21106","n_code_links":0,"syntology":null},{"paper":null,"title":"WebGuard++:Interpretable Malicious URL Detection via Bidirectional Fusion of HTML Subgraphs and Multi-Scale Convolutional BERT","date":"2025-06-24","arxiv_id":"2506.19356","n_code_links":0,"syntology":null},{"paper":null,"title":"Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning","date":"2025-06-12","arxiv_id":"2506.10629","n_code_links":0,"syntology":null},{"paper":"/paper/open-captchaworld-a-comprehensive-web-based","title":"Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents","date":"2025-05-30","arxiv_id":"2505.24878","n_code_links":1,"syntology":null},{"paper":null,"title":"MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection","date":"2025-05-26","arxiv_id":"2505.23803","n_code_links":0,"syntology":null},{"paper":null,"title":"URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training","date":"2025-05-22","arxiv_id":"2505.16570","n_code_links":0,"syntology":null},{"paper":"/paper/learning-better-representations-for-crowded","title":"Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection","date":"2025-05-21","arxiv_id":"2505.16029","n_code_links":1,"syntology":null},{"paper":"/paper/streamline-without-sacrifice-squeeze-out","title":"Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM","date":"2025-05-21","arxiv_id":"2505.15816","n_code_links":1,"syntology":{"ran":5,"of":9,"unverified":4,"pointer_only":4}},{"paper":"/paper/do-not-change-me-on-transferring-entities","title":"Do Not Change Me: On Transferring Entities Without Modification in Neural Machine Translation -- a Multilingual Perspective","date":"2025-05-09","arxiv_id":"2505.06010","n_code_links":2,"syntology":null},{"paper":"/paper/detecting-quishing-attacks-with-machine","title":"Detecting Quishing Attacks with Machine Learning Techniques Through QR Code Analysis","date":"2025-05-06","arxiv_id":"2505.03451","n_code_links":1,"syntology":null},{"paper":null,"title":"Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning","date":"2025-05-06","arxiv_id":"2505.03703","n_code_links":0,"syntology":null},{"paper":null,"title":"Helping Large Language Models Protect Themselves: An Enhanced Filtering and Summarization System","date":"2025-05-02","arxiv_id":"2505.01315","n_code_links":0,"syntology":null},{"paper":null,"title":"Phishing URL Detection using Bi-LSTM","date":"2025-04-29","arxiv_id":"2504.21049","n_code_links":0,"syntology":null},{"paper":null,"title":"A Gradient-Optimized TSK Fuzzy Framework for Explainable Phishing Detection","date":"2025-04-25","arxiv_id":"2504.18636","n_code_links":0,"syntology":null},{"paper":null,"title":"From Past to Present: A Survey of Malicious URL Detection Techniques, Datasets and Code Repositories","date":"2025-04-23","arxiv_id":"2504.16449","n_code_links":0,"syntology":null},{"paper":null,"title":"Emerging Cyber Attack Risks of Medical AI Agents","date":"2025-04-02","arxiv_id":"2504.03759","n_code_links":0,"syntology":null},{"paper":"/paper/mooncast-high-quality-zero-shot-podcast","title":"MoonCast: High-Quality Zero-Shot Podcast Generation","date":"2025-03-18","arxiv_id":"2503.14345","n_code_links":1,"syntology":null},{"paper":"/paper/visualwebinstruct-scaling-up-multimodal","title":"VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search","date":"2025-03-13","arxiv_id":"2503.10582","n_code_links":1,"syntology":{"ran":0,"of":13,"unverified":13,"pointer_only":0}},{"paper":"/paper/dongbamie-a-multimodal-information-extraction","title":"DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms","date":"2025-03-05","arxiv_id":"2503.03644","n_code_links":1,"syntology":null},{"paper":null,"title":"PhishVQC: Optimizing Phishing URL Detection with Correlation Based Feature Selection and Variational Quantum Classifier","date":"2025-03-03","arxiv_id":"2503.01799","n_code_links":0,"syntology":null},{"paper":"/paper/large-engagement-networks-for-classifying","title":"Large Engagement Networks for Classifying Coordinated Campaigns and Organic Twitter Trends","date":"2025-03-01","arxiv_id":"2503.00599","n_code_links":1,"syntology":null},{"paper":"/paper/tag-pag-a-dedicated-tool-for-systematic-web","title":"Tag-Pag: A Dedicated Tool for Systematic Web Page Annotations","date":"2025-02-22","arxiv_id":"2502.16150","n_code_links":1,"syntology":null},{"paper":null,"title":"CyberSentinel: An Emergent Threat Detection System for AI Security","date":"2025-02-20","arxiv_id":"2502.14966","n_code_links":0,"syntology":null},{"paper":"/paper/llm-agents-making-agent-tools","title":"LLM Agents Making Agent Tools","date":"2025-02-17","arxiv_id":"2502.11705","n_code_links":1,"syntology":null},{"paper":null,"title":"Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search","date":"2025-02-07","arxiv_id":"2502.04951","n_code_links":0,"syntology":null},{"paper":"/paper/a-new-dataset-and-methodology-for-malicious","title":"A New Dataset and Methodology for Malicious URL Classification","date":"2024-12-31","arxiv_id":"2501.00356","n_code_links":1,"syntology":null},{"paper":"/paper/gdsg-graph-diffusion-based-solution","title":"GDSG: Graph Diffusion-based Solution Generator for Optimization Problems in MEC Networks","date":"2024-12-11","arxiv_id":"2412.08296","n_code_links":1,"syntology":null},{"paper":null,"title":"Domain-specific Question Answering with Hybrid Search","date":"2024-12-04","arxiv_id":"2412.03736","n_code_links":0,"syntology":null},{"paper":"/paper/stain-aware-domain-alignment-for-imbalance","title":"Stain-aware Domain Alignment for Imbalance Blood Cell Classification","date":"2024-12-04","arxiv_id":"2412.02976","n_code_links":1,"syntology":null},{"paper":null,"title":"Large Multimodal Agents for Accurate Phishing Detection with Enhanced Token Optimization and Cost Reduction","date":"2024-12-03","arxiv_id":"2412.02301","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":3},{"task":"/task/benchmarking","name":"Benchmarking","papers":2},{"task":"/task/classification-1","name":"Classification","papers":2},{"task":"/task/zero-shot-learning","name":"Zero-Shot Learning","papers":2},{"task":null,"name":"zero-shot-classification","papers":2},{"task":"/task/3d-pedestrian-tracking","name":"3D Pedestrian Tracking","papers":1},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":1},{"task":"/task/binary-classification","name":"Binary Classification","papers":1},{"task":"/task/blocking","name":"Blocking","papers":1},{"task":"/task/code-generation","name":"Code Generation","papers":1},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/disentanglement","name":"Disentanglement","papers":1},{"task":"/task/diversity","name":"Diversity","papers":1},{"task":"/task/efficient-exploration","name":"Efficient Exploration","papers":1},{"task":"/task/feature-importance","name":"Feature Importance","papers":1},{"task":"/task/graph-classification","name":"Graph Classification","papers":1},{"task":"/task/graph-neural-network","name":"Graph Neural Network","papers":1},{"task":"/task/image-retrieval","name":"Image Retrieval","papers":1},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":1}],"tasks_shown":20,"n_tasks":50,"usage_by_year":[{"year":"2024","papers":9},{"year":"2025","papers":25}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/url"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}