{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-the-reward-function-for-a","title":"Learning the Reward Function for a Misspecified Model","arxiv_id":"1801.09624","date":"2018-01-29","proceeding":"ICML 2018 7","authors":["Erik Talvitie"],"abstract":"In model-based reinforcement learning it is typical to decouple the problems\nof learning the dynamics model and learning the reward function. However, when\nthe dynamics model is flawed, it may generate erroneous states that would never\noccur in the true environment. It is not clear a priori what value the reward\nfunction should assign to such states. This paper presents a novel error bound\nthat accounts for the reward model's behavior in states sampled from the model.\nThis bound is used to extend the existing Hallucinated DAgger-MC algorithm,\nwhich offers theoretical performance guarantees in deterministic MDPs that do\nnot assume a perfect model can be learned. Empirically, this approach to reward\nlearning can yield dramatic improvements in control performance when the\ndynamics model is flawed.","url_abs":"http://arxiv.org/abs/1801.09624v3","url_pdf":"http://arxiv.org/pdf/1801.09624v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-the-reward-function-for-a","repo_url":"https://github.com/etalvitie/hdaggermc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"model-based-reinforcement-learning","task_name":"Model-based Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.09624","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}