{"url":"/task/efficient-exploration","name":"Efficient Exploration","slug":"efficient-exploration","description_markdown":"**Efficient Exploration** is one of the main obstacles in scaling up modern deep reinforcement learning algorithms. The main challenge in Efficient Exploration is the balance between exploiting current estimates, and gaining information about poorly understood states and actions.\n\n\n<span class=\"description-source\">Source: [Randomized Value Functions via Multiplicative Normalizing Flows ](https://arxiv.org/abs/1806.02315)</span>","categories":[{"name":"Methodology","url":"/area/methodology"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":514,"papers_with_code":189,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":3,"subtasks":0,"parent_tasks":0},"benchmarks":[],"datasets":[{"url":"/dataset/replica","name":"Replica","full_name":"","num_papers_in_archive":414},{"url":"/dataset/house3d-environment","name":"House3D Environment","full_name":"","num_papers_in_archive":11},{"url":"/dataset/qdsd","name":"QDSD","full_name":"Quantum Dots Stability Diagrams","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":189,"tagged_in_all":514,"items":[{"url":"/paper/noisy-networks-for-exploration","title":"Noisy Networks for Exploration","date":"2017-06-30","arxiv_id":"1706.10295","repositories_listed":15,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":3}},{"url":"/paper/automatic-chemical-design-using-a-data-driven","title":"Automatic chemical design using a data-driven continuous representation of molecules","date":"2016-10-07","arxiv_id":"1610.02415","repositories_listed":11,"syntology":{"n":26,"n_ran":5,"n_unverified":21,"n_pointer_only":7}},{"url":"/paper/efficient-off-policy-meta-reinforcement","title":"Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables","date":"2019-03-19","arxiv_id":"1903.08254","repositories_listed":7,"syntology":{"n":9,"n_ran":6,"n_unverified":3,"n_pointer_only":1}},{"url":"/paper/stochastic-gradient-hamiltonian-monte-carlo","title":"Stochastic Gradient Hamiltonian Monte Carlo","date":"2014-02-17","arxiv_id":"1402.4102","repositories_listed":7,"syntology":{"n":6,"n_ran":0,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/deep-exploration-via-bootstrapped-dqn","title":"Deep Exploration via Bootstrapped DQN","date":"2016-02-15","arxiv_id":"1602.04621","repositories_listed":6,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/neural-contextual-bandits-with-upper","title":"Neural Contextual Bandits with UCB-based Exploration","date":"2019-11-11","arxiv_id":"1911.04462","repositories_listed":4,"syntology":null},{"url":"/paper/data-efficient-exploration-optimization-and","title":"Data-Efficient Exploration, Optimization, and Modeling of Diverse Designs through Surrogate-Assisted Illumination","date":"2017-02-13","arxiv_id":"1702.03713","repositories_listed":4,"syntology":null},{"url":"/paper/hyperagent-a-simple-scalable-efficient-and","title":"Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent","date":"2024-02-05","arxiv_id":"2402.10228","repositories_listed":3,"syntology":{"n":7,"n_ran":3,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/online-limited-memory-neural-linear-bandits-1","title":"Online Limited Memory Neural-Linear Bandits with Likelihood Matching","date":"2021-02-07","arxiv_id":"2102.03799","repositories_listed":3,"syntology":{"n":6,"n_ran":1,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/shared-experience-actor-critic-for-multi","title":"Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning","date":"2020-06-12","arxiv_id":"2006.07169","repositories_listed":3,"syntology":null},{"url":"/paper/scaling-map-elites-to-deep-neuroevolution","title":"Scaling MAP-Elites to Deep Neuroevolution","date":"2020-03-03","arxiv_id":"2003.01825","repositories_listed":3,"syntology":null},{"url":"/paper/conex-efficient-exploration-of-big-data","title":"ConEx: Efficient Exploration of Big-Data System Configurations for Better Performance","date":"2019-10-17","arxiv_id":"1910.09644","repositories_listed":3,"syntology":null},{"url":"/paper/scheduled-policy-optimization-for-natural","title":"Scheduled Policy Optimization for Natural Language Communication with Intelligent Agents","date":"2018-06-16","arxiv_id":"1806.06187","repositories_listed":3,"syntology":null},{"url":"/paper/adaptive-foundation-models-for-online","title":"Scalable Exploration via Ensemble++","date":"2024-07-18","arxiv_id":"2407.13195","repositories_listed":2,"syntology":{"n":8,"n_ran":5,"n_unverified":3,"n_pointer_only":8}},{"url":"/paper/streamlining-ocean-dynamics-modeling-with","title":"Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach","date":"2024-04-07","arxiv_id":"2404.05768","repositories_listed":2,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/hierarchical-spatial-proximity-reasoning-for","title":"Hierarchical Spatial Proximity Reasoning for Vision-and-Language Navigation","date":"2024-03-18","arxiv_id":"2403.11541","repositories_listed":2,"syntology":null},{"url":"/paper/demonstration-guided-reinforcement-learning-1","title":"Demonstration-Guided Reinforcement Learning with Efficient Exploration for Task Automation of Surgical Robot","date":"2023-02-20","arxiv_id":"2302.09772","repositories_listed":2,"syntology":null},{"url":"/paper/online-decision-transformer","title":"Online Decision Transformer","date":"2022-02-11","arxiv_id":"2202.05607","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/episodic-multi-agent-reinforcement-learning-1","title":"Episodic Multi-agent Reinforcement Learning with Curiosity-Driven Exploration","date":"2021-11-22","arxiv_id":"2111.11032","repositories_listed":2,"syntology":null},{"url":"/paper/paradiseo-from-a-modular-framework-for","title":"Paradiseo: From a Modular Framework for Evolutionary Computation to the Automated Design of Metaheuristics ---22 Years of Paradiseo---","date":"2021-05-02","arxiv_id":"2105.00420","repositories_listed":2,"syntology":null},{"url":"/paper/state-entropy-maximization-with-random","title":"State Entropy Maximization with Random Encoders for Efficient Exploration","date":"2021-02-18","arxiv_id":"2102.09430","repositories_listed":2,"syntology":null},{"url":"/paper/bebold-exploration-beyond-the-boundary-of-1","title":"BeBold: Exploration Beyond the Boundary of Explored Regions","date":"2020-12-15","arxiv_id":"2012.08621","repositories_listed":2,"syntology":null},{"url":"/paper/hybrid-genetic-search-for-the-cvrp-open","title":"Hybrid Genetic Search for the CVRP: Open-Source Implementation and SWAP* Neighborhood","date":"2020-11-23","arxiv_id":"2012.10384","repositories_listed":2,"syntology":null},{"url":"/paper/self-supervised-exploration-via-disagreement","title":"Self-Supervised Exploration via Disagreement","date":"2019-06-10","arxiv_id":"1906.04161","repositories_listed":2,"syntology":null},{"url":"/paper/estimating-risk-and-uncertainty-in-deep","title":"Estimating Risk and Uncertainty in Deep Reinforcement Learning","date":"2019-05-23","arxiv_id":"1905.09638","repositories_listed":2,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/learning-exploration-policies-for-navigation","title":"Learning Exploration Policies for Navigation","date":"2019-03-05","arxiv_id":"1903.01959","repositories_listed":2,"syntology":null},{"url":"/paper/model-based-active-exploration","title":"Model-Based Active Exploration","date":"2018-10-29","arxiv_id":"1810.12162","repositories_listed":2,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/nsga-net-a-multi-objective-genetic-algorithm","title":"NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm","date":"2018-10-08","arxiv_id":"1810.03522","repositories_listed":2,"syntology":null},{"url":"/paper/count-based-exploration-with-the-successor","title":"Count-Based Exploration with the Successor Representation","date":"2018-07-31","arxiv_id":"1807.11622","repositories_listed":2,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/randomized-value-functions-via-multiplicative","title":"Randomized Value Functions via Multiplicative Normalizing Flows","date":"2018-06-06","arxiv_id":"1806.02315","repositories_listed":2,"syntology":null}],"syntology_records":13,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}