{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-contextual-bandits-in-a-non","title":"Learning Contextual Bandits in a Non-stationary Environment","arxiv_id":"1805.09365","date":"2018-05-23","proceeding":null,"authors":["Qingyun Wu","Naveen Iyer","Hongning Wang"],"abstract":"Multi-armed bandit algorithms have become a reference solution for handling\nthe explore/exploit dilemma in recommender systems, and many other important\nreal-world problems, such as display advertisement. However, such algorithms\nusually assume a stationary reward distribution, which hardly holds in practice\nas users' preferences are dynamic. This inevitably costs a recommender system\nconsistent suboptimal performance. In this paper, we consider the situation\nwhere the underlying distribution of reward remains unchanged over (possibly\nshort) epochs and shifts at unknown time instants. In accordance, we propose a\ncontextual bandit algorithm that detects possible changes of environment based\non its reward estimation confidence and updates its arm selection strategy\nrespectively. Rigorous upper regret bound analysis of the proposed algorithm\ndemonstrates its learning effectiveness in such a non-trivial environment.\nExtensive empirical evaluations on both synthetic and real-world datasets for\nrecommendation confirm its practical utility in a changing environment.","url_abs":"http://arxiv.org/abs/1805.09365v1","url_pdf":"http://arxiv.org/pdf/1805.09365v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-contextual-bandits-in-a-non","repo_url":"https://github.com/YRussac/WeightedLinearBandits","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"multi-armed-bandits","task_name":"Multi-Armed Bandits"},{"task_slug":"recommendation-systems","task_name":"Recommendation Systems"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}