{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/marginal-policy-gradients-a-unified-family-of","title":"Marginal Policy Gradients: A Unified Family of Estimators for Bounded Action Spaces with Applications","arxiv_id":"1806.05134","date":"2018-06-13","proceeding":"ICLR 2019 5","authors":["Carson Eisenach","Haichuan Yang","Ji Liu","Han Liu"],"abstract":"Many complex domains, such as robotics control and real-time strategy (RTS)\ngames, require an agent to learn a continuous control. In the former, an agent\nlearns a policy over $\\mathbb{R}^d$ and in the latter, over a discrete set of\nactions each of which is parametrized by a continuous parameter. Such problems\nare naturally solved using policy based reinforcement learning (RL) methods,\nbut unfortunately these often suffer from high variance leading to instability\nand slow convergence. Unnecessary variance is introduced whenever policies over\nbounded action spaces are modeled using distributions with unbounded support by\napplying a transformation $T$ to the sampled action before execution in the\nenvironment. Recently, the variance reduced clipped action policy gradient\n(CAPG) was introduced for actions in bounded intervals, but to date no variance\nreduced methods exist when the action is a direction, something often seen in\nRTS games. To this end we introduce the angular policy gradient (APG), a\nstochastic policy gradient method for directional control. With the marginal\npolicy gradients family of estimators we present a unified analysis of the\nvariance reduction properties of APG and CAPG; our results provide a stronger\nguarantee than existing analyses for CAPG. Experimental results on a popular\nRTS game and a navigation task show that the APG estimator offers a substantial\nimprovement over the standard policy gradient.","url_abs":"http://arxiv.org/abs/1806.05134v3","url_pdf":"http://arxiv.org/pdf/1806.05134v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"marginal-policy-gradients-a-unified-family-of","repo_url":"https://github.com/ceisenach/MPG","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.05134","atlas_url":"https://app.syntology.ai/?focus=1806.05134","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}