{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/novelty-search-for-deep-reinforcement","title":"Novelty Search for Deep Reinforcement Learning Policy Network Weights by Action Sequence Edit Metric Distance","arxiv_id":"1902.03142","date":"2019-02-08","proceeding":null,"authors":["Ethan C. Jackson","Mark Daley"],"abstract":"Reinforcement learning (RL) problems often feature deceptive local optima,\nand learning methods that optimize purely for reward signal often fail to learn\nstrategies for overcoming them. Deep neuroevolution and novelty search have\nbeen proposed as effective alternatives to gradient-based methods for learning\nRL policies directly from pixels. In this paper, we introduce and evaluate the\nuse of novelty search over agent action sequences by string edit metric\ndistance as a means for promoting innovation. We also introduce a method for\nstagnation detection and population resampling inspired by recent developments\nin the RL community that uses the same mechanisms as novelty search to promote\nand develop innovative policies. Our methods extend a state-of-the-art method\nfor deep neuroevolution using a simple-yet-effective genetic algorithm (GA)\ndesigned to efficiently learn deep RL policy network weights. Experiments using\nfour games from the Atari 2600 benchmark were conducted. Results provide\nfurther evidence that GAs are competitive with gradient-based algorithms for\ndeep RL. Results also demonstrate that novelty search over action sequences is\nan effective source of selection pressure that can be integrated into existing\nevolutionary algorithms for deep RL.","url_abs":"http://arxiv.org/abs/1902.03142v1","url_pdf":"http://arxiv.org/pdf/1902.03142v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"novelty-search-for-deep-reinforcement","repo_url":"https://github.com/ethancjackson/NoveltySearchLevenshtein","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"evolutionary-algorithms","task_name":"Evolutionary Algorithms"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1902.03142","atlas_url":"https://app.syntology.ai/?focus=1902.03142","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}