{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/trust-pcl-an-off-policy-trust-region-method","title":"Trust-PCL: An Off-Policy Trust Region Method for Continuous Control","arxiv_id":"1707.01891","date":"2017-07-06","proceeding":"ICLR 2018 1","authors":["Ofir Nachum","Mohammad Norouzi","Kelvin Xu","Dale Schuurmans"],"abstract":"Trust region methods, such as TRPO, are often used to stabilize policy\noptimization algorithms in reinforcement learning (RL). While current trust\nregion strategies are effective for continuous control, they typically require\na prohibitively large amount of on-policy interaction with the environment. To\naddress this problem, we propose an off-policy trust region method, Trust-PCL.\nThe algorithm is the result of observing that the optimal policy and state\nvalues of a maximum reward objective with a relative-entropy regularizer\nsatisfy a set of multi-step pathwise consistencies along any path. Thus,\nTrust-PCL is able to maintain optimization stability while exploiting\noff-policy data to improve sample efficiency. When evaluated on a number of\ncontinuous control tasks, Trust-PCL improves the solution quality and sample\nefficiency of TRPO.","url_abs":"http://arxiv.org/abs/1707.01891v3","url_pdf":"http://arxiv.org/pdf/1707.01891v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"trust-pcl-an-off-policy-trust-region-method","repo_url":"https://github.com/tensorflow/models","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[{"method_slug":"trpo","method_name":"TRPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1707.01891","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}