{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/universal-reinforcement-learning-algorithms","title":"Universal Reinforcement Learning Algorithms: Survey and Experiments","arxiv_id":"1705.10557","date":"2017-05-30","proceeding":null,"authors":["John Aslanides","Jan Leike","Marcus Hutter"],"abstract":"Many state-of-the-art reinforcement learning (RL) algorithms typically assume\nthat the environment is an ergodic Markov Decision Process (MDP). In contrast,\nthe field of universal reinforcement learning (URL) is concerned with\nalgorithms that make as few assumptions as possible about the environment. The\nuniversal Bayesian agent AIXI and a family of related URL algorithms have been\ndeveloped in this setting. While numerous theoretical optimality results have\nbeen proven for these agents, there has been no empirical investigation of\ntheir behavior to date. We present a short and accessible survey of these URL\nalgorithms under a unified notation and framework, along with results of some\nexperiments that qualitatively illustrate some properties of the resulting\npolicies, and their relative performance on partially-observable gridworld\nenvironments. We also present an open-source reference implementation of the\nalgorithms which we hope will facilitate further understanding of, and\nexperimentation with, these ideas.","url_abs":"http://arxiv.org/abs/1705.10557v1","url_pdf":"http://arxiv.org/pdf/1705.10557v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"universal-reinforcement-learning-algorithms","repo_url":"https://github.com/aslanides/aixijs","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"survey","task_name":"Survey"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}