{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evolution-guided-policy-gradient-in","title":"Evolution-Guided Policy Gradient in Reinforcement Learning","arxiv_id":"1805.07917","date":"2018-05-21","proceeding":"NeurIPS 2018 12","authors":["Shauharda Khadka","Kagan Tumer"],"abstract":"Deep Reinforcement Learning (DRL) algorithms have been successfully applied\nto a range of challenging control tasks. However, these methods typically\nsuffer from three core difficulties: temporal credit assignment with sparse\nrewards, lack of effective exploration, and brittle convergence properties that\nare extremely sensitive to hyperparameters. Collectively, these challenges\nseverely limit the applicability of these approaches to real-world problems.\nEvolutionary Algorithms (EAs), a class of black box optimization techniques\ninspired by natural evolution, are well suited to address each of these three\nchallenges. However, EAs typically suffer from high sample complexity and\nstruggle to solve problems that require optimization of a large number of\nparameters. In this paper, we introduce Evolutionary Reinforcement Learning\n(ERL), a hybrid algorithm that leverages the population of an EA to provide\ndiversified data to train an RL agent, and reinserts the RL agent into the EA\npopulation periodically to inject gradient information into the EA. ERL\ninherits EA's ability of temporal credit assignment with a fitness metric,\neffective exploration with a diverse set of policies, and stability of a\npopulation-based approach and complements it with off-policy DRL's ability to\nleverage gradients for higher sample efficiency and faster learning.\nExperiments in a range of challenging continuous control benchmarks demonstrate\nthat ERL significantly outperforms prior DRL and EA methods.","url_abs":"http://arxiv.org/abs/1805.07917v2","url_pdf":"http://arxiv.org/pdf/1805.07917v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evolution-guided-policy-gradient-in","repo_url":"https://github.com/ShawK91/erl_paper_nips18","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"evolution-guided-policy-gradient-in","repo_url":"https://github.com/apourchot/ERL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"evolution-guided-policy-gradient-in","repo_url":"https://github.com/apourchot/ERL-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"evolution-guided-policy-gradient-in","repo_url":"https://github.com/apourchot/EvoLearn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"evolution-guided-policy-gradient-in","repo_url":"https://github.com/lyp741/Pytorch-ERL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"evolution-guided-policy-gradient-in","repo_url":"https://github.com/neilsgp/RL-Algorithms","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"evolutionary-algorithms","task_name":"Evolutionary Algorithms"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.07917","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.07917"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ShawK91/erl_paper_nips18","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lyp741/Pytorch-ERL","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/apourchot/ERL","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/apourchot/ERL-pytorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/apourchot/EvoLearn","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/neilsgp/RL-Algorithms","reach":{"status":"ok"}}],"summary":{"ran_fixture":1,"unverified":2},"by_repo_kind":{"listed":{"samples":3,"ran":1,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"825ecb6732e9fc31","entry":"fanin_init","repo":"lyp741/Pytorch-ERL","repo_kind":"listed","path":"pytorch-erl.py","file_url":"https://github.com/lyp741/Pytorch-ERL/blob/HEAD/pytorch-erl.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"825ecb6732e9fc31"}},{"code_sha256_prefix":"c788860ce99c0fc9","entry":"evaluate","repo":"apourchot/ERL-pytorch","repo_kind":"listed","path":"ERL.py","file_url":"https://github.com/apourchot/ERL-pytorch/blob/HEAD/ERL.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c788860ce99c0fc9"}},{"code_sha256_prefix":"5b08ad9b60a75183","entry":"train_rl","repo":"apourchot/ERL-pytorch","repo_kind":"listed","path":"ERL.py","file_url":"https://github.com/apourchot/ERL-pytorch/blob/HEAD/ERL.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5b08ad9b60a75183"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}