{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/periodic-bandits-and-wireless-network","title":"Periodic Bandits and Wireless Network Selection","arxiv_id":"1904.12355","date":"2019-04-28","proceeding":null,"authors":["Shunhao Oh","Anuja Meetoo Appavoo","Seth Gilbert"],"abstract":"Bandit-style algorithms have been studied extensively in stochastic and\nadversarial settings. Such algorithms have been shown to be useful in\nmultiplayer settings, e.g. to solve the wireless network selection problem,\nwhich can be formulated as an adversarial bandit problem. A leading bandit\nalgorithm for the adversarial setting is EXP3. However, network behavior is\noften repetitive, where user density and network behavior follow regular\npatterns. Bandit algorithms, like EXP3, fail to provide good guarantees for\nperiodic behaviors. A major reason is that these algorithms compete against\nfixed-action policies, which is ineffective in a periodic setting.\n  In this paper, we define a periodic bandit setting, and periodic regret as a\nbetter performance measure for this type of setting. Instead of comparing an\nalgorithm's performance to fixed-action policies, we aim to be competitive with\npolicies that play arms under some set of possible periodic patterns $F$ (for\nexample, all possible periodic functions with periods $1,2,\\cdots,P$). We\npropose Periodic EXP4, a computationally efficient variant of the EXP4\nalgorithm for periodic settings. With $K$ arms, $T$ time steps, and where each\nperiodic pattern in $F$ is of length at most $P$, we show that the periodic\nregret obtained by Periodic EXP4 is at most $O\\big(\\sqrt{PKT \\log K + KT \\log\n|F|}\\big)$. We also prove a lower bound of $\\Omega\\big(\\sqrt{PKT + KT\n\\frac{\\log |F|}{\\log K}} \\big)$ for the periodic setting, showing that this is\noptimal within log-factors. As an example, we focus on the wireless network\nselection problem. Through simulation, we show that Periodic EXP4 learns the\nperiodic pattern over time, adapts to changes in a dynamic environment, and far\noutperforms EXP3.","url_abs":"http://arxiv.org/abs/1904.12355v1","url_pdf":"http://arxiv.org/pdf/1904.12355v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"periodic-bandits-and-wireless-network","repo_url":"https://github.com/Ohohcakester/PeriodicEXP4-Source","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}