{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/guide-actor-critic-for-continuous-control","title":"Guide Actor-Critic for Continuous Control","arxiv_id":"1705.07606","date":"2017-05-22","proceeding":"ICLR 2018 1","authors":["Voot Tangkaratt","Abbas Abdolmaleki","Masashi Sugiyama"],"abstract":"Actor-critic methods solve reinforcement learning problems by updating a\nparameterized policy known as an actor in a direction that increases an\nestimate of the expected return known as a critic. However, existing\nactor-critic methods only use values or gradients of the critic to update the\npolicy parameter. In this paper, we propose a novel actor-critic method called\nthe guide actor-critic (GAC). GAC firstly learns a guide actor that locally\nmaximizes the critic and then it updates the policy parameter based on the\nguide actor by supervised learning. Our main theoretical contributions are two\nfolds. First, we show that GAC updates the guide actor by performing\nsecond-order optimization in the action space where the curvature matrix is\nbased on the Hessians of the critic. Second, we show that the deterministic\npolicy gradient method is a special case of GAC when the Hessians are ignored.\nThrough experiments, we show that our method is a promising reinforcement\nlearning method for continuous controls.","url_abs":"http://arxiv.org/abs/1705.07606v2","url_pdf":"http://arxiv.org/pdf/1705.07606v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"guide-actor-critic-for-continuous-control","repo_url":"https://github.com/voot-t/guide-actor-critic","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.07606","atlas_url":"https://app.syntology.ai/?focus=1705.07606","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}