{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/metacontrol-for-adaptive-imagination-based","title":"Metacontrol for Adaptive Imagination-Based Optimization","arxiv_id":"1705.02670","date":"2017-05-07","proceeding":null,"authors":["Jessica B. Hamrick","Andrew J. Ballard","Razvan Pascanu","Oriol Vinyals","Nicolas Heess","Peter W. Battaglia"],"abstract":"Many machine learning systems are built to solve the hardest examples of a\nparticular task, which often makes them large and expensive to run---especially\nwith respect to the easier examples, which might require much less computation.\nFor an agent with a limited computational budget, this \"one-size-fits-all\"\napproach may result in the agent wasting valuable computation on easy examples,\nwhile not spending enough on hard examples. Rather than learning a single,\nfixed policy for solving all instances of a task, we introduce a metacontroller\nwhich learns to optimize a sequence of \"imagined\" internal simulations over\npredictive models of the world in order to construct a more informed, and more\neconomical, solution. The metacontroller component is a model-free\nreinforcement learning agent, which decides both how many iterations of the\noptimization procedure to run, as well as which model to consult on each\niteration. The models (which we call \"experts\") can be state transition models,\naction-value functions, or any other mechanism that provides information useful\nfor solving the task, and can be learned on-policy or off-policy in parallel\nwith the metacontroller. When the metacontroller, controller, and experts were\ntrained with \"interaction networks\" (Battaglia et al., 2016) as expert models,\nour approach was able to solve a challenging decision-making problem under\ncomplex non-linear dynamics. The metacontroller learned to adapt the amount of\ncomputation it performed to the difficulty of the task, and learned how to\nchoose which experts to consult by factoring in both their reliability and\nindividual computational resource costs. This allowed the metacontroller to\nachieve a lower overall cost (task loss plus computational cost) than more\ntraditional fixed policy approaches. These results demonstrate that our\napproach is a powerful framework for using...","url_abs":"http://arxiv.org/abs/1705.02670v1","url_pdf":"http://arxiv.org/pdf/1705.02670v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"metacontrol-for-adaptive-imagination-based","repo_url":"https://github.com/deepmind/spaceship_dataset","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[{"slug":"spaceship-dataset","name":"Spaceship Dataset","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.02670","atlas_url":"https://app.syntology.ai/?focus=1705.02670","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}