{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rotting-infinitely-many-armed-bandits","title":"Rotting Infinitely Many-armed Bandits","arxiv_id":"2201.12975","date":"2022-01-31","proceeding":null,"authors":["Jung-hun Kim","Milan Vojnovic","Se-Young Yun"],"abstract":"We consider the infinitely many-armed bandit problem with rotting rewards, where the mean reward of an arm decreases at each pull of the arm according to an arbitrary trend with maximum rotting rate $\\varrho=o(1)$. We show that this learning problem has an $\\Omega(\\max\\{\\varrho^{1/3}T,\\sqrt{T}\\})$ worst-case regret lower bound where $T$ is the horizon time. We show that a matching upper bound $\\tilde{O}(\\max\\{\\varrho^{1/3}T,\\sqrt{T}\\})$, up to a poly-logarithmic factor, can be achieved by an algorithm that uses a UCB index for each arm and a threshold value to decide whether to continue pulling an arm or remove the arm from further consideration, when the algorithm knows the value of the maximum rotting rate $\\varrho$. We also show that an $\\tilde{O}(\\max\\{\\varrho^{1/3}T,T^{3/4}\\})$ regret upper bound can be achieved by an algorithm that does not know the value of $\\varrho$, by using an adaptive UCB index along with an adaptive threshold value.","url_abs":"https://arxiv.org/abs/2201.12975v3","url_pdf":"https://arxiv.org/pdf/2201.12975v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rotting-infinitely-many-armed-bandits","repo_url":"https://github.com/junghunkim7786/rotting_infinite_armed_bandits","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2201.12975","atlas_url":"https://app.syntology.ai/?focus=2201.12975","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}