{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/td-regularized-actor-critic-methods","title":"TD-Regularized Actor-Critic Methods","arxiv_id":"1812.08288","date":"2018-12-19","proceeding":null,"authors":["Simone Parisi","Voot Tangkaratt","Jan Peters","Mohammad Emtiyaz Khan"],"abstract":"Actor-critic methods can achieve incredible performance on difficult\nreinforcement learning problems, but they are also prone to instability. This\nis partly due to the interaction between the actor and critic during learning,\ne.g., an inaccurate step taken by one of them might adversely affect the other\nand destabilize the learning. To avoid such issues, we propose to regularize\nthe learning objective of the actor by penalizing the temporal difference (TD)\nerror of the critic. This improves stability by avoiding large steps in the\nactor update whenever the critic is highly inaccurate. The resulting method,\nwhich we call the TD-regularized actor-critic method, is a simple plug-and-play\napproach to improve stability and overall performance of the actor-critic\nmethods. Evaluations on standard benchmarks confirm this.","url_abs":"http://arxiv.org/abs/1812.08288v3","url_pdf":"http://arxiv.org/pdf/1812.08288v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"td-regularized-actor-critic-methods","repo_url":"https://github.com/sparisi/td-reg","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.08288","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}