{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stacked-thompson-bandits","title":"Stacked Thompson Bandits","arxiv_id":"1702.08726","date":"2017-02-28","proceeding":null,"authors":["Lenz Belzner","Thomas Gabor"],"abstract":"We introduce Stacked Thompson Bandits (STB) for efficiently generating plans\nthat are likely to satisfy a given bounded temporal logic requirement. STB uses\na simulation for evaluation of plans, and takes a Bayesian approach to using\nthe resulting information to guide its search. In particular, we show that\nstacking multiarmed bandits and using Thompson sampling to guide the action\nselection process for each bandit enables STB to generate plans that satisfy\nrequirements with a high probability while only searching a fraction of the\nsearch space.","url_abs":"http://arxiv.org/abs/1702.08726v1","url_pdf":"http://arxiv.org/pdf/1702.08726v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stacked-thompson-bandits","repo_url":"https://github.com/jazzbob/stb","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"thompson-sampling","task_name":"Thompson Sampling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}