{"url":"/task/multi-armed-bandits","name":"Multi-Armed Bandits","slug":"multi-armed-bandits","description_markdown":"Multi-armed bandits refer to a task where a fixed amount of resources must be allocated between competing resources that maximizes expected gain. Typically these problems involve an exploration/exploitation trade-off.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Microsoft Research](http://research.microsoft.com/en-us/projects/bandits/) )</span>","categories":[{"name":"Methodology","url":"/area/methodology"},{"name":"Miscellaneous","url":"/area/miscellaneous"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":1262,"papers_with_code":253,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":2,"subtasks":1,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/multi-armed-bandits-on-mushroom","slug":"multi-armed-bandits-on-mushroom","dataset":"Mushroom","dataset_url":null,"rows_in_archive":2,"metrics":["Cumulative regret"],"first_row_in_archive_order":{"model":"Linear FullPosterior-MR","paper_title":"Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling","paper_url":"/paper/deep-bayesian-bandits-showdown-an-empirical","paper_date":"2018-02-26","arxiv_id":"1802.09127","code_links":[{"title":"tensorflow/models","url":"https://github.com/tensorflow/models"},{"title":"tensorflow/models","url":"https://github.com/tensorflow/models/tree/archive/research/deep_contextual_bandits"},{"title":"mlisicki/neuralkernelbandits","url":"https://github.com/mlisicki/neuralkernelbandits"},{"title":"vectorinstitute/neuralkernelbandits","url":"https://github.com/vectorinstitute/neuralkernelbandits"}],"syntology":null}}],"datasets":[{"url":"/dataset/kuairand","name":"KuaiRand","full_name":"","num_papers_in_archive":42},{"url":"/dataset/duolingo-notifications-data","name":"Duolingo Bandit Notifications","full_name":"","num_papers_in_archive":1}],"subtasks":[{"url":"/task/thompson-sampling","name":"Thompson Sampling"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":253,"tagged_in_all":1262,"items":[{"url":"/paper/deep-reinforcement-learning-based","title":"Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions Modeling","date":"2018-10-29","arxiv_id":"1810.12027","repositories_listed":5,"syntology":null},{"url":"/paper/neural-contextual-bandits-with-upper","title":"Neural Contextual Bandits with UCB-based Exploration","date":"2019-11-11","arxiv_id":"1911.04462","repositories_listed":4,"syntology":null},{"url":"/paper/deep-bayesian-bandits-showdown-an-empirical","title":"Deep Bayesian Bandits Showdown: An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling","date":"2018-02-26","arxiv_id":"1802.09127","repositories_listed":4,"syntology":null},{"url":"/paper/hypothesis-generation-with-large-language","title":"Hypothesis Generation with Large Language Models","date":"2024-04-05","arxiv_id":"2404.04326","repositories_listed":3,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/off-policy-evaluation-for-large-action-spaces","title":"Off-Policy Evaluation for Large Action Spaces via Embeddings","date":"2022-02-13","arxiv_id":"2202.06317","repositories_listed":3,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/online-limited-memory-neural-linear-bandits-1","title":"Online Limited Memory Neural-Linear Bandits with Likelihood Matching","date":"2021-02-07","arxiv_id":"2102.03799","repositories_listed":3,"syntology":{"n":6,"n_ran":1,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/gaussian-gated-linear-networks","title":"Gaussian Gated Linear Networks","date":"2020-06-10","arxiv_id":"2006.05964","repositories_listed":3,"syntology":null},{"url":"/paper/locally-differentially-private-contextual","title":"Locally Differentially Private (Contextual) Bandits Learning","date":"2020-06-01","arxiv_id":"2006.00701","repositories_listed":3,"syntology":null},{"url":"/paper/on-line-adaptative-curriculum-learning-for","title":"On-line Adaptative Curriculum Learning for GANs","date":"2018-07-31","arxiv_id":"1808.00020","repositories_listed":3,"syntology":null},{"url":"/paper/optimal-and-adaptive-off-policy-evaluation-in","title":"Optimal and Adaptive Off-policy Evaluation in Contextual Bandits","date":"2016-12-04","arxiv_id":"1612.01205","repositories_listed":3,"syntology":null},{"url":"/paper/adaptive-foundation-models-for-online","title":"Scalable Exploration via Ensemble++","date":"2024-07-18","arxiv_id":"2407.13195","repositories_listed":2,"syntology":{"n":8,"n_ran":5,"n_unverified":3,"n_pointer_only":8}},{"url":"/paper/optimal-regret-with-limited-adaptivity-for","title":"Generalized Linear Bandits with Limited Adaptivity","date":"2024-04-10","arxiv_id":"2404.06831","repositories_listed":2,"syntology":{"n":11,"n_ran":10,"n_unverified":1,"n_pointer_only":4}},{"url":"/paper/an-experimental-design-for-anytime-valid","title":"An Experimental Design for Anytime-Valid Causal Inference on Multi-Armed Bandits","date":"2023-11-09","arxiv_id":"2311.05794","repositories_listed":2,"syntology":null},{"url":"/paper/kernel-conditional-moment-constraints-for","title":"Kernel Conditional Moment Constraints for Confounding Robust Inference","date":"2023-02-26","arxiv_id":"2302.13348","repositories_listed":2,"syntology":{"n":17,"n_ran":0,"n_unverified":17,"n_pointer_only":0}},{"url":"/paper/truncated-linucb-for-stochastic-linear","title":"Truncated LinUCB for Stochastic Linear Bandits","date":"2022-02-23","arxiv_id":"2202.11735","repositories_listed":2,"syntology":null},{"url":"/paper/doubly-robust-off-policy-evaluation-for","title":"Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model","date":"2022-02-03","arxiv_id":"2202.01562","repositories_listed":2,"syntology":null},{"url":"/paper/an-empirical-study-of-neural-kernel-bandits","title":"Empirical analysis of representation learning and exploration in neural kernel bandits","date":"2021-11-05","arxiv_id":"2111.03543","repositories_listed":2,"syntology":null},{"url":"/paper/inverse-contextual-bandits-learning-how","title":"Inverse Contextual Bandits: Learning How Behavior Evolves over Time","date":"2021-07-13","arxiv_id":"2107.06317","repositories_listed":2,"syntology":null},{"url":"/paper/quantile-bandits-for-best-arms-identification","title":"Quantile Bandits for Best Arms Identification","date":"2020-10-22","arxiv_id":"2010.11568","repositories_listed":2,"syntology":null},{"url":"/paper/neural-thompson-sampling","title":"Neural Thompson Sampling","date":"2020-10-02","arxiv_id":"2010.00827","repositories_listed":2,"syntology":{"n":7,"n_ran":6,"n_unverified":1,"n_pointer_only":7}},{"url":"/paper/dual-mandate-patrols-multi-armed-bandits-for","title":"Dual-Mandate Patrols: Multi-Armed Bandits for Green Security","date":"2020-09-14","arxiv_id":"2009.06560","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/confident-off-policy-evaluation-and-selection","title":"Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting","date":"2020-06-18","arxiv_id":"2006.10460","repositories_listed":2,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/bandit-pam-almost-linear-time-k-medoids","title":"BanditPAM: Almost Linear Time $k$-Medoids Clustering via Multi-Armed Bandits","date":"2020-06-11","arxiv_id":"2006.06856","repositories_listed":2,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":5}},{"url":"/paper/optimal-and-greedy-algorithms-for-multi-armed","title":"The Unreasonable Effectiveness of Greedy Algorithms in Multi-Armed Bandit with Many Arms","date":"2020-02-24","arxiv_id":"2002.10121","repositories_listed":2,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/a-modern-introduction-to-online-learning","title":"A Modern Introduction to Online Learning","date":"2019-12-31","arxiv_id":"1912.13213","repositories_listed":2,"syntology":null},{"url":"/paper/multi-armed-bandits-with-correlated-arms","title":"Multi-Armed Bandits with Correlated Arms","date":"2019-11-06","arxiv_id":"1911.03959","repositories_listed":2,"syntology":null},{"url":"/paper/bayesian-optimisation-over-multiple","title":"Bayesian Optimisation over Multiple Continuous and Categorical Inputs","date":"2019-06-20","arxiv_id":"1906.08878","repositories_listed":2,"syntology":{"n":7,"n_ran":0,"n_unverified":7,"n_pointer_only":0}},{"url":"/paper/intrinsically-efficient-stable-and-bounded","title":"Intrinsically Efficient, Stable, and Bounded Off-Policy Evaluation for Reinforcement Learning","date":"2019-06-09","arxiv_id":"1906.03735","repositories_listed":2,"syntology":null},{"url":"/paper/adapting-multi-armed-bandits-policies-to","title":"Adapting multi-armed bandits policies to contextual bandits scenarios","date":"2018-11-11","arxiv_id":"1811.04383","repositories_listed":2,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/sic-mmab-synchronisation-involves","title":"SIC-MMAB: Synchronisation Involves Communication in Multiplayer Multi-Armed Bandits","date":"2018-09-21","arxiv_id":"1809.08151","repositories_listed":2,"syntology":null}],"syntology_records":13,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}