{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bayesian-bandits-balancing-the-exploration","title":"Bayesian bandits: balancing the exploration-exploitation tradeoff via double sampling","arxiv_id":"1709.03162","date":"2017-09-10","proceeding":null,"authors":["Iñigo Urteaga","Chris H. Wiggins"],"abstract":"Reinforcement learning studies how to balance exploration and exploitation in\nreal-world systems, optimizing interactions with the world while simultaneously\nlearning how the world operates. One general class of algorithms for such\nlearning is the multi-armed bandit setting. Randomized probability matching,\nbased upon the Thompson sampling approach introduced in the 1930s, has recently\nbeen shown to perform well and to enjoy provable optimality properties. It\npermits generative, interpretable modeling in a Bayesian setting, where prior\nknowledge is incorporated, and the computed posteriors naturally capture the\nfull state of knowledge. In this work, we harness the information contained in\nthe Bayesian posterior and estimate its sufficient statistics via sampling. In\nseveral application domains, for example in health and medicine, each\ninteraction with the world can be expensive and invasive, whereas drawing\nsamples from the model is relatively inexpensive. Exploiting this viewpoint, we\ndevelop a double sampling technique driven by the uncertainty in the learning\nprocess: it favors exploitation when certain about the properties of each arm,\nexploring otherwise. The proposed algorithm does not make any distributional\nassumption and it is applicable to complex reward distributions, as long as\nBayesian posterior updates are computable. Utilizing the estimated posterior\nsufficient statistics, double sampling autonomously balances the\nexploration-exploitation tradeoff to make better informed decisions. We\nempirically show its reduced cumulative regret when compared to\nstate-of-the-art alternatives in representative bandit settings.","url_abs":"http://arxiv.org/abs/1709.03162v2","url_pdf":"http://arxiv.org/pdf/1709.03162v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bayesian-bandits-balancing-the-exploration","repo_url":"https://github.com/iurteaga/bandits","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"thompson-sampling","task_name":"Thompson Sampling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1709.03162","atlas_url":"https://app.syntology.ai/?focus=1709.03162","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}