{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalised-discount-functions-applied-to-a","title":"Generalised Discount Functions applied to a Monte-Carlo AImu Implementation","arxiv_id":"1703.01358","date":"2017-03-03","proceeding":null,"authors":["Sean Lamont","John Aslanides","Jan Leike","Marcus Hutter"],"abstract":"In recent years, work has been done to develop the theory of General\nReinforcement Learning (GRL). However, there are few examples demonstrating\nthese results in a concrete way. In particular, there are no examples\ndemonstrating the known results regarding gener- alised discounting. We have\nadded to the GRL simulation platform AIXIjs the functionality to assign an\nagent arbitrary discount functions, and an environment which can be used to\ndetermine the effect of discounting on an agent's policy. Using this, we\ninvestigate how geometric, hyperbolic and power discounting affect an informed\nagent in a simple MDP. We experimentally reproduce a number of theoretical\nresults, and discuss some related subtleties. It was found that the agent's\nbehaviour followed what is expected theoretically, assuming appropriate\nparameters were chosen for the Monte-Carlo Tree Search (MCTS) planning\nalgorithm.","url_abs":"http://arxiv.org/abs/1703.01358v1","url_pdf":"http://arxiv.org/pdf/1703.01358v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalised-discount-functions-applied-to-a","repo_url":"https://github.com/aslanides/aixijs","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"general-reinforcement-learning","task_name":"General Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"monte-carlo-tree-search","method_name":"Monte-Carlo Tree Search"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}