{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-bias-span-constrained-exploration","title":"Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning","arxiv_id":"1802.04020","date":"2018-02-12","proceeding":"ICML 2018 7","authors":["Ronan Fruit","Matteo Pirotta","Alessandro Lazaric","Ronald Ortner"],"abstract":"We introduce SCAL, an algorithm designed to perform efficient\nexploration-exploitation in any unknown weakly-communicating Markov decision\nprocess (MDP) for which an upper bound $c$ on the span of the optimal bias\nfunction is known. For an MDP with $S$ states, $A$ actions and $\\Gamma \\leq S$\npossible next states, we prove a regret bound of $\\widetilde{O}(c\\sqrt{\\Gamma\nSAT})$, which significantly improves over existing algorithms (e.g., UCRL and\nPSRL), whose regret scales linearly with the MDP diameter $D$. In fact, the\noptimal bias span is finite and often much smaller than $D$ (e.g., $D=\\infty$\nin non-communicating MDPs). A similar result was originally derived by Bartlett\nand Tewari (2009) for REGAL.C, for which no tractable algorithm is available.\nIn this paper, we relax the optimization problem at the core of REGAL.C, we\ncarefully analyze its properties, and we provide the first computationally\nefficient algorithm to solve it. Finally, we report numerical simulations\nsupporting our theoretical findings and showing how SCAL significantly\noutperforms UCRL in MDPs with large diameter and small span.","url_abs":"http://arxiv.org/abs/1802.04020v2","url_pdf":"http://arxiv.org/pdf/1802.04020v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-bias-span-constrained-exploration","repo_url":"https://github.com/RonanFR/UCRL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"efficient-exploration","task_name":"Efficient Exploration"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.04020","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}