{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-exploration-in-evolution-strategies","title":"Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents","arxiv_id":"1712.06560","date":"2017-12-18","proceeding":"NeurIPS 2018 12","authors":["Edoardo Conti","Vashisht Madhavan","Felipe Petroski Such","Joel Lehman","Kenneth O. Stanley","Jeff Clune"],"abstract":"Evolution strategies (ES) are a family of black-box optimization algorithms\nable to train deep neural networks roughly as well as Q-learning and policy\ngradient methods on challenging deep reinforcement learning (RL) problems, but\nare much faster (e.g. hours vs. days) because they parallelize better. However,\nmany RL problems require directed exploration because they have reward\nfunctions that are sparse or deceptive (i.e. contain local optima), and it is\nunknown how to encourage such exploration with ES. Here we show that algorithms\nthat have been invented to promote directed exploration in small-scale evolved\nneural networks via populations of exploring agents, specifically novelty\nsearch (NS) and quality diversity (QD) algorithms, can be hybridized with ES to\nimprove its performance on sparse or deceptive deep RL tasks, while retaining\nscalability. Our experiments confirm that the resultant new algorithms, NS-ES\nand two QD algorithms, NSR-ES and NSRA-ES, avoid local optima encountered by ES\nto achieve higher performance on Atari and simulated robots learning to walk\naround a deceptive trap. This paper thus introduces a family of fast, scalable\nalgorithms for reinforcement learning that are capable of directed exploration.\nIt also adds this new family of exploration algorithms to the RL toolbox and\nraises the interesting possibility that analogous algorithms with multiple\nsimultaneous paths of exploration might also combine well with existing RL\nalgorithms outside ES.","url_abs":"http://arxiv.org/abs/1712.06560v3","url_pdf":"http://arxiv.org/pdf/1712.06560v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-exploration-in-evolution-strategies","repo_url":"https://github.com/uber-research/deep-neuroevolution","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"improving-exploration-in-evolution-strategies","repo_url":"https://github.com/uber-common/deep-neuroevolution","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"policy-gradient-methods","task_name":"Policy Gradient Methods"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.06560","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1712.06560"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/uber-common/deep-neuroevolution","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/uber-research/deep-neuroevolution","reach":{"status":"unanswered"}}],"summary":{"ran_violates":1,"ran_fixture":1,"unverified":1},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"c8c2960f6687b93c","entry":"compute_centered_ranks","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"c8c2960f6687b93c"}},{"code_sha256_prefix":"e2d29b359d6b9487","entry":"compute_ranks","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"e2d29b359d6b9487"}},{"code_sha256_prefix":"703807555aa85c6d","entry":"make_session","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"703807555aa85c6d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}