{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ex2-exploration-with-exemplar-models-for-deep","title":"EX2: Exploration with Exemplar Models for Deep Reinforcement Learning","arxiv_id":"1703.01260","date":"2017-03-03","proceeding":"NeurIPS 2017 12","authors":["Justin Fu","John D. Co-Reyes","Sergey Levine"],"abstract":"Deep reinforcement learning algorithms have been shown to learn complex tasks\nusing highly general policy classes. However, sparse reward problems remain a\nsignificant challenge. Exploration methods based on novelty detection have been\nparticularly successful in such settings but typically require generative or\npredictive models of the observations, which can be difficult to train when the\nobservations are very high-dimensional and complex, as in the case of raw\nimages. We propose a novelty detection algorithm for exploration that is based\nentirely on discriminatively trained exemplar models, where classifiers are\ntrained to discriminate each visited state against all others. Intuitively,\nnovel states are easier to distinguish against other states seen during\ntraining. We show that this kind of discriminative modeling corresponds to\nimplicit density estimation, and that it can be combined with count-based\nexploration to produce competitive results on a range of popular benchmark\ntasks, including state-of-the-art results on challenging egocentric\nobservations in the vizDoom benchmark.","url_abs":"http://arxiv.org/abs/1703.01260v2","url_pdf":"http://arxiv.org/pdf/1703.01260v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ex2-exploration-with-exemplar-models-for-deep","repo_url":"https://github.com/jcoreyes/ex2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"density-estimation","task_name":"Density Estimation"},{"task_slug":"novelty-detection","task_name":"Novelty Detection"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.01260","atlas_url":"https://app.syntology.ai/?focus=1703.01260","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}