{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pipps-flexible-model-based-policy-search","title":"PIPPS: Flexible Model-Based Policy Search Robust to the Curse of Chaos","arxiv_id":"1902.01240","date":"2019-02-04","proceeding":"ICML 2018 7","authors":["Paavo Parmas","Carl Edward Rasmussen","Jan Peters","Kenji Doya"],"abstract":"Previously, the exploding gradient problem has been explained to be central\nin deep learning and model-based reinforcement learning, because it causes\nnumerical issues and instability in optimization. Our experiments in\nmodel-based reinforcement learning imply that the problem is not just a\nnumerical issue, but it may be caused by a fundamental chaos-like nature of\nlong chains of nonlinear computations. Not only do the magnitudes of the\ngradients become large, the direction of the gradients becomes essentially\nrandom. We show that reparameterization gradients suffer from the problem,\nwhile likelihood ratio gradients are robust. Using our insights, we develop a\nmodel-based policy search framework, Probabilistic Inference for Particle-Based\nPolicy Search (PIPPS), which is easily extensible, and allows for almost\narbitrary models and policies, while simultaneously matching the performance of\nprevious data-efficient learning algorithms. Finally, we invent the total\npropagation algorithm, which efficiently computes a union over all pathwise\nderivative depths during a single backwards pass, automatically giving greater\nweight to estimators with lower variance, sometimes improving over\nreparameterization gradients by $10^6$ times.","url_abs":"http://arxiv.org/abs/1902.01240v1","url_pdf":"http://arxiv.org/pdf/1902.01240v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pipps-flexible-model-based-policy-search","repo_url":"https://github.com/proppo/pipps_demo","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"pipps-flexible-model-based-policy-search","repo_url":"https://github.com/Archibald-Lafraik/pipps","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok"}},{"paper_slug":"pipps-flexible-model-based-policy-search","repo_url":"https://github.com/natolambert/dynamicslearn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"model-based-reinforcement-learning","task_name":"Model-based Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.01240","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}