Browse State-of-the-Art › Multi-Armed Bandits
Multi-Armed Bandits
253 papers with code · 1 benchmark · 2 datasets archive 2025-07-28
Multi-armed bandits refer to a task where a fixed amount of resources must be allocated between competing resources that maximizes expected gain. Typically these problems involve an exploration/exploitation trade-off.
( Image credit: Microsoft Research )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Mushroom (2 rows) | Linear FullPosterior-MR | Deep Bayesian Bandits Showdown: An Empirical Comparison of... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 253 papers with code (1,262 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Oct 2018 5 repositories listedThe DRR framework treats recommendation as a sequential decision making procedure and adopts an "Actor-Critic" reinforcement learning scheme to model the interactions between the users and recommender systems, which can…
-
11 Nov 2019 4 repositories listedTo the best of our knowledge, it is the first neural network-based contextual bandit algorithm with a near-optimal regret guarantee.
-
26 Feb 2018 4 repositories listedAt the same time, advances in approximate Bayesian methods have made posterior approximation for flexible neural network models practical.
-
5 Apr 2024 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We focus on hypothesis generation based on data (i.
-
13 Feb 2022 3 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedUnfortunately, when the number of actions is large, existing OPE estimators -- most of which are based on inverse propensity score weighting -- degrade severely and can suffer from extreme bias and variance.
-
7 Feb 2021 3 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedTo alleviate this, we propose a likelihood matching algorithm that is resilient to catastrophic forgetting and is completely online.
-
10 Jun 2020 3 repositories listedWe propose the Gaussian Gated Linear Network (G-GLN), an extension to the recently proposed GLN family of deep neural networks.
-
1 Jun 2020 3 repositories listedWe study locally differentially private (LDP) bandits learning in this paper.
-
31 Jul 2018 3 repositories listedWe argue that less expressive discriminators are smoother and have a general coarse grained view of the modes map, which enforces the generator to cover a wide portion of the data distribution support.
-
4 Dec 2016 3 repositories listedWe study the off-policy evaluation problem---estimating the value of a target policy using data collected by another policy---under the contextual bandit model.
-
18 Jul 2024 2 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)Thompson Sampling is a principled method for balancing exploration and exploitation, but its real-world adoption faces computational challenges in large-scale or non-conjugate settings.
-
10 Apr 2024 2 repositories listed Syntology ran 10 of 11 samples · 1 unverified · 4 pointer-only (licence)We study the generalized linear contextual bandit problem within the constraints of limited adaptivity.
-
9 Nov 2023 2 repositories listedThis paper addresses both challenges by introducing the Mixture Adaptive Design (MAD), a new experimental design for multi-armed bandit (MAB) algorithms that enables anytime-valid inference on the Average Treatment…
-
26 Feb 2023 2 repositories listed Syntology ran 0 of 17 samples · 17 unverifiedIt can be shown that our estimator contains the recently proposed sharp estimator by Dorn and Guo (2022) as a special case, and our method enables a novel extension of the classical marginal sensitivity model using…
-
23 Feb 2022 2 repositories listedThis paper considers contextual bandits with a finite number of arms, where the contexts are independent and identically distributed d-dimensional random vectors, and the expected rewards are linear in both the arm…
-
3 Feb 2022 2 repositories listedWe show that the proposed estimator is unbiased in more cases compared to existing estimators that make stronger assumptions.
-
5 Nov 2021 2 repositories listedWe consider policies based on a GP and a Student's t-process (TP).
-
13 Jul 2021 2 repositories listedUnderstanding a decision-maker's priorities by observing their behavior is critical for transparency and accountability in decision processes, such as in healthcare.
-
22 Oct 2020 2 repositories listedWe consider a variant of the best arm identification task in stochastic multi-armed bandits.
-
2 Oct 2020 2 repositories listed Syntology ran 6 of 7 samples · 1 unverified · 7 pointer-only (licence)Thompson Sampling (TS) is one of the most effective algorithms for solving contextual multi-armed bandit problems.
-
14 Sep 2020 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedConservation efforts in green security domains to protect wildlife and forests are constrained by the limited availability of defenders (i.
-
18 Jun 2020 2 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedWe consider off-policy evaluation in the contextual bandit setting for the purpose of obtaining a robust off-policy selection strategy, where the selection strategy is evaluated based on the value of the chosen policy…
-
11 Jun 2020 2 repositories listed Syntology ran 0 of 5 samples · 5 unverified · 5 pointer-only (licence)Current state-of-the-art k-medoids clustering algorithms, such as Partitioning Around Medoids (PAM), are iterative and are quadratic in the dataset size n for each iteration, being prohibitively expensive for large…
-
24 Feb 2020 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThis finding diverges from the notion of free exploration, which relates to covariate variation, as recently discussed in contextual bandit literature.
-
31 Dec 2019 2 repositories listedI present first-order and second-order algorithms for online learning with convex losses, in Euclidean and non-Euclidean settings.
-
6 Nov 2019 2 repositories listedWe consider a multi-armed bandit framework where the rewards obtained by pulling different arms are correlated.
-
20 Jun 2019 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedEfficient optimisation of black-box problems that comprise both continuous and categorical inputs is important, yet poses significant challenges.
-
9 Jun 2019 2 repositories listedWe propose new estimators for OPE based on empirical likelihood that are always more efficient than IS, SNIS, and DR and satisfy the same stability and boundedness properties as SNIS.
-
11 Nov 2018 2 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedThis work explores adaptations of successful multi-armed bandits policies to the online contextual bandits scenario with binary rewards using binary classification algorithms such as logistic regression as black-box…
-
21 Sep 2018 2 repositories listedMotivated by cognitive radio networks, we consider the stochastic multiplayer multi-armed bandit problem, where several players pull arms simultaneously and collisions occur if one of them is pulled by several players…
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections