{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/model-learning-for-look-ahead-exploration-in","title":"Model Learning for Look-ahead Exploration in Continuous Control","arxiv_id":"1811.08086","date":"2018-11-20","proceeding":null,"authors":["Arpit Agarwal","Katharina Muelling","Katerina Fragkiadaki"],"abstract":"We propose an exploration method that incorporates look-ahead search over\nbasic learnt skills and their dynamics, and use it for reinforcement learning\n(RL) of manipulation policies . Our skills are multi-goal policies learned in\nisolation in simpler environments using existing multigoal RL formulations,\nanalogous to options or macroactions. Coarse skill dynamics, i.e., the state\ntransition caused by a (complete) skill execution, are learnt and are unrolled\nforward during lookahead search. Policy search benefits from temporal\nabstraction during exploration, though itself operates over low-level primitive\nactions, and thus the resulting policies does not suffer from suboptimality and\ninflexibility caused by coarse skill chaining. We show that the proposed\nexploration strategy results in effective learning of complex manipulation\npolicies faster than current state-of-the-art RL methods, and converges to\nbetter policies than methods that use options or parametrized skills as\nbuilding blocks of the policy itself, as opposed to guiding exploration. We\nshow that the proposed exploration strategy results in effective learning of\ncomplex manipulation policies faster than current state-of-the-art RL methods,\nand converges to better policies than methods that use options or parameterized\nskills as building blocks of the policy itself, as opposed to guiding\nexploration.","url_abs":"http://arxiv.org/abs/1811.08086v1","url_pdf":"http://arxiv.org/pdf/1811.08086v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"model-learning-for-look-ahead-exploration-in","repo_url":"https://github.com/arpit15/skill-based-exploration-drl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.08086","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}