{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/guided-cost-learning-deep-inverse-optimal","title":"Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization","arxiv_id":"1603.00448","date":"2016-03-01","proceeding":null,"authors":["Chelsea Finn","Sergey Levine","Pieter Abbeel"],"abstract":"Reinforcement learning can acquire complex behaviors from high-level\nspecifications. However, defining a cost function that can be optimized\neffectively and encodes the correct task is challenging in practice. We explore\nhow inverse optimal control (IOC) can be used to learn behaviors from\ndemonstrations, with applications to torque control of high-dimensional robotic\nsystems. Our method addresses two key challenges in inverse optimal control:\nfirst, the need for informative features and effective regularization to impose\nstructure on the cost, and second, the difficulty of learning the cost function\nunder unknown dynamics for high-dimensional continuous systems. To address the\nformer challenge, we present an algorithm capable of learning arbitrary\nnonlinear cost functions, such as neural networks, without meticulous feature\nengineering. To address the latter challenge, we formulate an efficient\nsample-based approximation for MaxEnt IOC. We evaluate our method on a series\nof simulated tasks and real-world robotic manipulation problems, demonstrating\nsubstantial improvement over prior methods both in terms of task complexity and\nsample efficiency.","url_abs":"http://arxiv.org/abs/1603.00448v3","url_pdf":"http://arxiv.org/pdf/1603.00448v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"guided-cost-learning-deep-inverse-optimal","repo_url":"https://github.com/Alina-Samokhina/guided_cost_RL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"guided-cost-learning-deep-inverse-optimal","repo_url":"https://github.com/ninawie/guided-cost-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"guided-cost-learning-deep-inverse-optimal","repo_url":"https://github.com/nishantkr18/guided-cost-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"guided-cost-learning-deep-inverse-optimal","repo_url":"https://github.com/opendilab/DI-engine","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"feature-engineering","task_name":"Feature Engineering"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1603.00448","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}