{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/guided-policy-search-as-approximate-mirror","title":"Guided Policy Search as Approximate Mirror Descent","arxiv_id":"1607.04614","date":"2016-07-15","proceeding":null,"authors":["William Montgomery","Sergey Levine"],"abstract":"Guided policy search algorithms can be used to optimize complex nonlinear\npolicies, such as deep neural networks, without directly computing policy\ngradients in the high-dimensional parameter space. Instead, these methods use\nsupervised learning to train the policy to mimic a \"teacher\" algorithm, such as\na trajectory optimizer or a trajectory-centric reinforcement learning method.\nGuided policy search methods provide asymptotic local convergence guarantees by\nconstruction, but it is not clear how much the policy improves within a small,\nfinite number of iterations. We show that guided policy search algorithms can\nbe interpreted as an approximate variant of mirror descent, where the\nprojection onto the constraint manifold is not exact. We derive a new guided\npolicy search algorithm that is simpler and provides appealing improvement and\nconvergence guarantees in simplified convex and linear settings, and show that\nin the more general nonlinear setting, the error in the projection step can be\nbounded. We provide empirical results on several simulated robotic navigation\nand manipulation tasks that show that our method is stable and achieves similar\nor better performance when compared to prior guided policy search methods, with\na simpler formulation and fewer hyperparameters.","url_abs":"http://arxiv.org/abs/1607.04614v1","url_pdf":"http://arxiv.org/pdf/1607.04614v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"guided-policy-search-as-approximate-mirror","repo_url":"https://github.com/cbfinn/gps","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1607.04614","atlas_url":"https://app.syntology.ai/?focus=1607.04614","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}