{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/utility-decomposition-with-deep-corrections","title":"Decomposition Methods with Deep Corrections for Reinforcement Learning","arxiv_id":"1802.01772","date":"2018-02-06","proceeding":null,"authors":["Maxime Bouton","Kyle Julian","Alireza Nakhaei","Kikuo Fujimura","Mykel J. Kochenderfer"],"abstract":"Decomposition methods have been proposed to approximate solutions to large\nsequential decision making problems. In contexts where an agent interacts with\nmultiple entities, utility decomposition can be used to separate the global\nobjective into local tasks considering each individual entity independently. An\narbitrator is then responsible for combining the individual utilities and\nselecting an action in real time to solve the global problem. Although these\ntechniques can perform well empirically, they rely on strong assumptions of\nindependence between the local tasks and sacrifice the optimality of the global\nsolution. This paper proposes an approach that improves upon such approximate\nsolutions by learning a correction term represented by a neural network. We\ndemonstrate this approach on a fisheries management problem where multiple\nboats must coordinate to maximize their catch over time as well as on a\npedestrian avoidance problem for autonomous driving. In each problem,\ndecomposition methods can scale to multiple boats or pedestrians by using\nstrategies involving one entity. We verify empirically that the proposed\ncorrection method significantly improves the decomposition method and\noutperforms a policy trained on the full scale problem without utility\ndecomposition.","url_abs":"http://arxiv.org/abs/1802.01772v2","url_pdf":"http://arxiv.org/pdf/1802.01772v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"utility-decomposition-with-deep-corrections","repo_url":"https://github.com/sisl/AutomotivePOMDPs.jl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"management","task_name":"Management"},{"task_slug":"problem-decomposition","task_name":"Problem Decomposition"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"sequential-decision-making","task_name":"Sequential Decision Making"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}