{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/near-optimal-exploration-exploitation-in-non","title":"Near Optimal Exploration-Exploitation in Non-Communicating Markov Decision Processes","arxiv_id":"1807.02373","date":"2018-07-06","proceeding":"NeurIPS 2018 12","authors":["Ronan Fruit","Matteo Pirotta","Alessandro Lazaric"],"abstract":"While designing the state space of an MDP, it is common to include states\nthat are transient or not reachable by any policy (e.g., in mountain car, the\nproduct space of speed and position contains configurations that are not\nphysically reachable). This leads to defining weakly-communicating or\nmulti-chain MDPs. In this paper, we introduce \\tucrl, the first algorithm able\nto perform efficient exploration-exploitation in any finite Markov Decision\nProcess (MDP) without requiring any form of prior knowledge. In particular, for\nany MDP with $S^{\\texttt{C}}$ communicating states, $A$ actions and\n$\\Gamma^{\\texttt{C}} \\leq S^{\\texttt{C}}$ possible communicating next states,\nwe derive a $\\widetilde{O}(D^{\\texttt{C}} \\sqrt{\\Gamma^{\\texttt{C}}\nS^{\\texttt{C}} AT})$ regret bound, where $D^{\\texttt{C}}$ is the diameter\n(i.e., the longest shortest path) of the communicating part of the MDP. This is\nin contrast with optimistic algorithms (e.g., UCRL, Optimistic PSRL) that\nsuffer linear regret in weakly-communicating MDPs, as well as posterior\nsampling or regularised algorithms (e.g., REGAL), which require prior knowledge\non the bias span of the optimal policy to bias the exploration to achieve\nsub-linear regret. We also prove that in weakly-communicating MDPs, no\nalgorithm can ever achieve a logarithmic growth of the regret without first\nsuffering a linear regret for a number of steps that is exponential in the\nparameters of the MDP. Finally, we report numerical simulations supporting our\ntheoretical findings and showing how TUCRL overcomes the limitations of the\nstate-of-the-art.","url_abs":"http://arxiv.org/abs/1807.02373v2","url_pdf":"http://arxiv.org/pdf/1807.02373v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"near-optimal-exploration-exploitation-in-non","repo_url":"https://github.com/RonanFR/UCRL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"efficient-exploration","task_name":"Efficient Exploration"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.02373","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}