{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/beyond-confidence-regions-tight-bayesian","title":"Beyond Confidence Regions: Tight Bayesian Ambiguity Sets for Robust MDPs","arxiv_id":"1902.07605","date":"2019-02-20","proceeding":"NeurIPS 2019 12","authors":["Marek Petrik","Reazul Hasan Russell"],"abstract":"Robust MDPs (RMDPs) can be used to compute policies with provable worst-case\nguarantees in reinforcement learning. The quality and robustness of an RMDP\nsolution are determined by the ambiguity set---the set of plausible transition\nprobabilities---which is usually constructed as a multi-dimensional confidence\nregion. Existing methods construct ambiguity sets as confidence regions using\nconcentration inequalities which leads to overly conservative solutions. This\npaper proposes a new paradigm that can achieve better solutions with the same\nrobustness guarantees without using confidence regions as ambiguity sets. To\nincorporate prior knowledge, our algorithms optimize the size and position of\nambiguity sets using Bayesian inference. Our theoretical analysis shows the\nsafety of the proposed method, and the empirical results demonstrate its\npractical promise.","url_abs":"http://arxiv.org/abs/1902.07605v1","url_pdf":"http://arxiv.org/pdf/1902.07605v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"beyond-confidence-regions-tight-bayesian","repo_url":"https://github.com/marekpetrik/craam2","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"bayesian-inference","task_name":"Bayesian Inference"},{"task_slug":null,"task_name":"Position"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.07605","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}