{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-regularization-for-markov-decision","title":"Temporal Regularization for Markov Decision Process","arxiv_id":null,"date":"2018-12-01","proceeding":"NeurIPS 2018 12","authors":["Pierre Thodoroff","Audrey Durand","Joelle Pineau","Doina Precup"],"abstract":"Several applications of Reinforcement Learning suffer from instability due to high\nvariance. This is especially prevalent in high dimensional domains. Regularization\nis a commonly used technique in machine learning to reduce variance, at the cost\nof introducing some bias. Most existing regularization techniques focus on spatial\n(perceptual) regularization. Yet in reinforcement learning, due to the nature of the\nBellman equation, there is an opportunity to also exploit temporal regularization\nbased on smoothness in value estimates over trajectories. This paper explores a\nclass of methods for temporal regularization. We formally characterize the bias\ninduced by this technique using Markov chain concepts. We illustrate the various\ncharacteristics of temporal regularization via a sequence of simple discrete and\ncontinuous MDPs, and show that the technique provides improvement even in\nhigh-dimensional Atari games.","url_abs":"http://papers.nips.cc/paper/7449-temporal-regularization-for-markov-decision-process","url_pdf":"http://papers.nips.cc/paper/7449-temporal-regularization-for-markov-decision-process.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporal-regularization-for-markov-decision","repo_url":"https://github.com/pierthodo/temporal_regularization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}