{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/corralling-a-band-of-bandit-algorithms","title":"Corralling a Band of Bandit Algorithms","arxiv_id":"1612.06246","date":"2016-12-19","proceeding":null,"authors":["Alekh Agarwal","Haipeng Luo","Behnam Neyshabur","Robert E. Schapire"],"abstract":"We study the problem of combining multiple bandit algorithms (that is, online\nlearning algorithms with partial feedback) with the goal of creating a master\nalgorithm that performs almost as well as the best base algorithm if it were to\nbe run on its own. The main challenge is that when run with a master, base\nalgorithms unavoidably receive much less feedback and it is thus critical that\nthe master not starve a base algorithm that might perform uncompetitively\ninitially but would eventually outperform others if given enough feedback. We\naddress this difficulty by devising a version of Online Mirror Descent with a\nspecial mirror map together with a sophisticated learning rate scheme. We show\nthat this approach manages to achieve a more delicate balance between\nexploiting and exploring base algorithms than previous works yielding superior\nregret bounds.\n  Our results are applicable to many settings, such as multi-armed bandits,\ncontextual bandits, and convex bandits. As examples, we present two main\napplications. The first is to create an algorithm that enjoys worst-case\nrobustness while at the same time performing much better when the environment\nis relatively easy. The second is to create an algorithm that works\nsimultaneously under different assumptions of the environment, such as\ndifferent priors or different loss structures.","url_abs":"http://arxiv.org/abs/1612.06246v3","url_pdf":"http://arxiv.org/pdf/1612.06246v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"corralling-a-band-of-bandit-algorithms","repo_url":"https://github.com/lasgroup/alexp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"multi-armed-bandits","task_name":"Multi-Armed Bandits"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1612.06246","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}