{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hyperbolic-discounting-and-learning-over","title":"Hyperbolic Discounting and Learning over Multiple Horizons","arxiv_id":"1902.06865","date":"2019-02-19","proceeding":"ICLR 2020 1","authors":["William Fedus","Carles Gelada","Yoshua Bengio","Marc G. Bellemare","Hugo Larochelle"],"abstract":"Reinforcement learning (RL) typically defines a discount factor as part of\nthe Markov Decision Process. The discount factor values future rewards by an\nexponential scheme that leads to theoretical convergence guarantees of the\nBellman equation. However, evidence from psychology, economics and neuroscience\nsuggests that humans and animals instead have hyperbolic time-preferences. In\nthis work we revisit the fundamentals of discounting in RL and bridge this\ndisconnect by implementing an RL agent that acts via hyperbolic discounting. We\ndemonstrate that a simple approach approximates hyperbolic discount functions\nwhile still using familiar temporal-difference learning techniques in RL.\nAdditionally, and independent of hyperbolic discounting, we make a surprising\ndiscovery that simultaneously learning value functions over multiple\ntime-horizons is an effective auxiliary task which often improves over a strong\nvalue-based RL agent, Rainbow.","url_abs":"http://arxiv.org/abs/1902.06865v3","url_pdf":"http://arxiv.org/pdf/1902.06865v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hyperbolic-discounting-and-learning-over","repo_url":"https://github.com/google-research/google-research/tree/master/hyperbolic_discount","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.06865","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}