{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-model-based-deep-reinforcement","title":"Efficient Model-Based Deep Reinforcement Learning with Variational State Tabulation","arxiv_id":"1802.04325","date":"2018-02-12","proceeding":"ICML 2018 7","authors":["Dane Corneil","Wulfram Gerstner","Johanni Brea"],"abstract":"Modern reinforcement learning algorithms reach super-human performance on\nmany board and video games, but they are sample inefficient, i.e. they\ntypically require significantly more playing experience than humans to reach an\nequal performance level. To improve sample efficiency, an agent may build a\nmodel of the environment and use planning methods to update its policy. In this\narticle we introduce Variational State Tabulation (VaST), which maps an\nenvironment with a high-dimensional state space (e.g. the space of visual\ninputs) to an abstract tabular model. Prioritized sweeping with small backups,\na highly efficient planning method, can then be used to update state-action\nvalues. We show how VaST can rapidly learn to maximize reward in tasks like 3D\nnavigation and efficiently adapt to sudden changes in rewards or transition\nprobabilities.","url_abs":"http://arxiv.org/abs/1802.04325v2","url_pdf":"http://arxiv.org/pdf/1802.04325v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-model-based-deep-reinforcement","repo_url":"https://github.com/danecor/VaST","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"prioritized-sweeping","method_name":"Prioritized Sweeping"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.04325","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}