{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/taming-the-noise-in-reinforcement-learning","title":"Taming the Noise in Reinforcement Learning via Soft Updates","arxiv_id":"1512.08562","date":"2015-12-28","proceeding":null,"authors":["Roy Fox","Ari Pakman","Naftali Tishby"],"abstract":"Model-free reinforcement learning algorithms, such as Q-learning, perform\npoorly in the early stages of learning in noisy environments, because much\neffort is spent unlearning biased estimates of the state-action value function.\nThe bias results from selecting, among several noisy estimates, the apparent\noptimum, which may actually be suboptimal. We propose G-learning, a new\noff-policy learning algorithm that regularizes the value estimates by\npenalizing deterministic policies in the beginning of the learning process. We\nshow that this method reduces the bias of the value-function estimation,\nleading to faster convergence to the optimal value and the optimal policy.\nMoreover, G-learning enables the natural incorporation of prior domain\nknowledge, when available. The stochastic nature of G-learning also makes it\navoid some exploration costs, a property usually attributed only to on-policy\nalgorithms. We illustrate these ideas in several examples, where G-learning\nresults in significant improvements of the convergence rate and the cost of the\nlearning process.","url_abs":"http://arxiv.org/abs/1512.08562v4","url_pdf":"http://arxiv.org/pdf/1512.08562v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"taming-the-noise-in-reinforcement-learning","repo_url":"https://github.com/jollyraven100/Quant_algorithmic-trading_and_More","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"taming-the-noise-in-reinforcement-learning","repo_url":"https://github.com/jollyraven100/Trade-Ideas-and-Research-Reference","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"taming-the-noise-in-reinforcement-learning","repo_url":"https://github.com/michaelsyao/Trade-Ideas-and-Research-Reference","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1512.08562","atlas_url":"https://app.syntology.ai/?focus=1512.08562","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}