{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/relative-entropy-regularized-policy-iteration","title":"Relative Entropy Regularized Policy Iteration","arxiv_id":"1812.02256","date":"2018-12-05","proceeding":null,"authors":["Abbas Abdolmaleki","Jost Tobias Springenberg","Jonas Degrave","Steven Bohez","Yuval Tassa","Dan Belov","Nicolas Heess","Martin Riedmiller"],"abstract":"We present an off-policy actor-critic algorithm for Reinforcement Learning\n(RL) that combines ideas from gradient-free optimization via stochastic search\nwith learned action-value function. The result is a simple procedure consisting\nof three steps: i) policy evaluation by estimating a parametric action-value\nfunction; ii) policy improvement via the estimation of a local non-parametric\npolicy; and iii) generalization by fitting a parametric policy. Each step can\nbe implemented in different ways, giving rise to several algorithm variants.\nOur algorithm draws on connections to existing literature on black-box\noptimization and 'RL as an inference' and it can be seen either as an extension\nof the Maximum a Posteriori Policy Optimisation algorithm (MPO) [Abdolmaleki et\nal., 2018a], or as an extension of Trust Region Covariance Matrix Adaptation\nEvolutionary Strategy (CMA-ES) [Abdolmaleki et al., 2017b; Hansen et al., 1997]\nto a policy iteration scheme. Our comparison on 31 continuous control tasks\nfrom parkour suite [Heess et al., 2017], DeepMind control suite [Tassa et al.,\n2018] and OpenAI Gym [Brockman et al., 2016] with diverse properties, limited\namount of compute and a single set of hyperparameters, demonstrate the\neffectiveness of our method and the state of art results. Videos, summarizing\nresults, can be found at goo.gl/HtvJKR .","url_abs":"http://arxiv.org/abs/1812.02256v1","url_pdf":"http://arxiv.org/pdf/1812.02256v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"relative-entropy-regularized-policy-iteration","repo_url":"https://github.com/acyclics/MPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"openai-gym","task_name":"OpenAI Gym"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.02256","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}