{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/model-based-planning-with-discrete-and","title":"Model-Based Planning with Discrete and Continuous Actions","arxiv_id":"1705.07177","date":"2017-05-19","proceeding":null,"authors":["Mikael Henaff","William F. Whitney","Yann Lecun"],"abstract":"Action planning using learned and differentiable forward models of the world\nis a general approach which has a number of desirable properties, including\nimproved sample complexity over model-free RL methods, reuse of learned models\nacross different tasks, and the ability to perform efficient gradient-based\noptimization in continuous action spaces. However, this approach does not apply\nstraightforwardly when the action space is discrete. In this work, we show that\nit is in fact possible to effectively perform planning via backprop in discrete\naction spaces, using a simple paramaterization of the actions vectors on the\nsimplex combined with input noise when training the forward model. Our\nexperiments show that this approach can match or outperform model-free RL and\ndiscrete planning methods on gridworld navigation tasks in terms of performance\nand/or planning time while using limited environment interactions, and can\nadditionally be used to perform model-based control in a challenging new task\nwhere the action space combines discrete and continuous actions. We furthermore\npropose a policy distillation approach which yields a fast policy network which\ncan be used at inference time, removing the need for an iterative planning\nprocedure.","url_abs":"http://arxiv.org/abs/1705.07177v2","url_pdf":"http://arxiv.org/pdf/1705.07177v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"model-based-planning-with-discrete-and","repo_url":"https://github.com/jackdawe/joliRL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":{"status":"unanswered"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.07177","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}