Browse State-of-the-Art › Multi-Armed Bandits › Papers, page 5
Multi-Armed Bandits
Papers archive 2025-07-28
archive papers tagged: 1,262 · with a code link: 253 · where Syntology ran a sample: 55 (44 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (55 of 1,262 tagged: 44 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 5 of 13: papers 401 to 500 of 1,262, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Empathic Responding for Digital Interpersonal Emotion Regulation via Content Recommendation5 Aug 2024 0 repositories listed
-
Online Learning for Autonomous Management of Intent-based 6G Networks25 Jul 2024 0 repositories listed
-
Identifiable latent bandits: Combining observational data and exploration for personalized healthcare23 Jul 2024 0 repositories listed
-
Satisficing Exploration for Deep Reinforcement Learning16 Jul 2024 0 repositories listed
-
Open Problem: Tight Bounds for Kernelized Multi-Armed Bandits with Bernoulli Rewards8 Jul 2024 0 repositories listed
-
Honor Among Bandits: No-Regret Learning for Online Fair Division1 Jul 2024 0 repositories listed
-
A Contextual Combinatorial Bandit Approach to Negotiation30 Jun 2024 0 repositories listed
-
Classical Bandit Algorithms for Entanglement Detection in Parameterized Qubit States28 Jun 2024 0 repositories listed
-
EduQate: Generating Adaptive Curricula through RMABs in Education Settings20 Jun 2024 0 repositories listed
-
BEACON: Balancing Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes19 Jun 2024 0 repositories listed
-
Towards Bayesian Data Selection18 Jun 2024 0 repositories listed
-
Improving Reward-Conditioned Policies for Multi-Armed Bandits using Normalized Weight Functions16 Jun 2024 0 repositories listed
-
An Adaptive Method for Contextual Stochastic Multi-armed Bandits with Rewards Generated by a Linear Dynamical System14 Jun 2024 0 repositories listed
-
Towards Domain Adaptive Neural Contextual Bandits13 Jun 2024 0 repositories listed
-
A Federated Online Restless Bandit Framework for Cooperative Resource Allocation12 Jun 2024 0 repositories listed
-
Asymptotically Optimal Regret for Black-Box Predict-then-Optimize12 Jun 2024 0 repositories listed
-
Sample Complexity Reduction via Policy Difference Estimation in Tabular Reinforcement Learning11 Jun 2024 0 repositories listed
-
A conversion theorem and minimax optimality for continuum contextual bandits9 Jun 2024 0 repositories listed
-
Data-Driven Upper Confidence Bounds with Near-Optimal Regret for Heavy-Tailed Bandits9 Jun 2024 0 repositories listed
-
Adaptively Learning to Select-Rank in Online Platforms7 Jun 2024 0 repositories listed
-
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond3 Jun 2024 0 repositories listed
-
Global Rewards in Restless Multi-Armed Bandits2 Jun 2024 0 repositories listed
-
A Batch Sequential Halving Algorithm without Performance Degradation1 Jun 2024 0 repositories listed
-
Strategic Linear Contextual Bandits1 Jun 2024 0 repositories listed
-
No-Regret Learning for Fair Multi-Agent Social Welfare Optimization31 May 2024 0 repositories listed
-
Understanding Memory-Regret Trade-Off for Streaming Stochastic Multi-Armed Bandits30 May 2024 0 repositories listed
-
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff28 May 2024 0 repositories listed
-
Optimizing Sharpe Ratio: Risk-Adjusted Decision-Making in Multi-Armed Bandits28 May 2024 0 repositories listed
-
Multi-Player Approaches for Dueling Bandits25 May 2024 0 repositories listed
-
Indexed Minimum Empirical Divergence-Based Algorithms for Linear Bandits24 May 2024 0 repositories listed
-
Budgeted Recommendation with Delayed Feedback19 May 2024 0 repositories listed
-
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization10 May 2024 0 repositories listed
-
Federated Combinatorial Multi-Agent Multi-Armed Bandits9 May 2024 0 repositories listed
-
Imprecise Multi-Armed Bandits9 May 2024 0 repositories listed
-
Leveraging (Biased) Information: Multi-armed Bandits with Offline Data4 May 2024 0 repositories listed
-
Mathematics of statistical sequential decision-making: concentration, risk-awareness and modelling in stochastic bandits, with applications to bariatric surgery3 May 2024 0 repositories listed
-
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback2 May 2024 0 repositories listed
-
Recommenadation aided Caching using Combinatorial Multi-armed Bandits30 Apr 2024 0 repositories listed
-
Disentangling Exploration from Exploitation29 Apr 2024 0 repositories listed
-
Structured Reinforcement Learning for Delay-Optimal Data Transmission in Dense mmWave Networks25 Apr 2024 0 repositories listed
-
Feel-Good Thompson Sampling for Contextual Dueling Bandits9 Apr 2024 0 repositories listed
-
On the Importance of Uncertainty in Decision-Making with Large Language Models3 Apr 2024 0 repositories listed
-
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy2 Apr 2024 0 repositories listed
-
Nearly-tight Approximation Guarantees for the Improving Multi-Armed Bandits Problem1 Apr 2024 0 repositories listed
-
A Correction of Pseudo Log-Likelihood Method26 Mar 2024 0 repositories listed
-
Contextual Restless Multi-Armed Bandits with Application to Demand Response Decision-Making22 Mar 2024 0 repositories listed
-
Transfer in Sequential Multi-armed Bandits via Reward Samples19 Mar 2024 0 repositories listed
-
Phasic Diversity Optimization for Population-Based Reinforcement Learning17 Mar 2024 0 repositories listed
-
ε-Neural Thompson Sampling of Deep Brain Stimulation for Parkinson Disease Treatment11 Mar 2024 0 repositories listed
-
Cramming Contextual Bandits for On-policy Statistical Evaluation11 Mar 2024 0 repositories listed
-
Efficient Public Health Intervention Planning Using Decomposition-Based Decision-Focused Learning8 Mar 2024 0 repositories listed
-
A General Reduction for High-Probability Analysis with General Light-Tailed Distributions5 Mar 2024 0 repositories listed
-
LC-Tsallis-INF: Generalized Best-of-Both-Worlds Linear Contextual Bandits5 Mar 2024 0 repositories listed
-
Adaptive Learning Rate for Follow-the-Regularized-Leader: Competitive Analysis and Best-of-Both-Worlds1 Mar 2024 0 repositories listed
-
Federated Linear Contextual Bandits with Heterogeneous Clients29 Feb 2024 0 repositories listed
-
Investigating Gender Fairness in Machine Learning-driven Personalized Care for Chronic Pain29 Feb 2024 0 repositories listed
-
Batched Nonparametric Contextual Bandits27 Feb 2024 0 repositories listed
-
Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancement24 Feb 2024 0 repositories listed
-
Multi-Armed Bandits with Abstention23 Feb 2024 0 repositories listed
-
Optimistic Information Directed Sampling23 Feb 2024 0 repositories listed
-
A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health22 Feb 2024 0 repositories listed
-
Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits21 Feb 2024 0 repositories listed
-
Incentivized Exploration via Filtered Posterior Sampling20 Feb 2024 0 repositories listed
-
Efficient Prompt Optimization Through the Lens of Best Arm Identification15 Feb 2024 0 repositories listed
-
Diffusion Models Meet Contextual Bandits with Large Action Spaces15 Feb 2024 0 repositories listed
-
Thompson Sampling in Partially Observable Contextual Bandits15 Feb 2024 0 repositories listed
-
FLASH: Federated Learning Across Simultaneous Heterogeneities13 Feb 2024 0 repositories listed
-
Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits13 Feb 2024 0 repositories listed
-
Contextual Multinomial Logit Bandits with General Value Functions12 Feb 2024 0 repositories listed
-
Efficient Contextual Bandits with Uninformed Feedback Graphs12 Feb 2024 0 repositories listed
-
Replicability is Asymptotically Free in Multi-armed Bandits12 Feb 2024 0 repositories listed
-
Stochastic contextual bandits with graph feedback: from independence number to MAS number12 Feb 2024 0 repositories listed
-
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning11 Feb 2024 0 repositories listed
-
Fast UCB-type algorithms for stochastic bandits with heavy and super heavy symmetric noise10 Feb 2024 0 repositories listed
-
Tree Ensembles for Contextual Bandits10 Feb 2024 0 repositories listed
-
Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits7 Feb 2024 0 repositories listed
-
Fairness and Privacy Guarantees in Federated Contextual Bandits5 Feb 2024 0 repositories listed
-
Multi-Armed Bandits with Interference2 Feb 2024 0 repositories listed
-
Query-Efficient Correlation Clustering with Noisy Oracle2 Feb 2024 0 repositories listed
-
Distributed Multi-Task Learning for Stochastic Bandits with Context Distribution and Stage-wise Constraints21 Jan 2024 0 repositories listed
-
Adaptive Regret for Bandits Made Possible: Two Queries Suffice17 Jan 2024 0 repositories listed
-
On Quantum Natural Policy Gradients16 Jan 2024 0 repositories listed
-
Contextual Bandits with Stage-wise Constraints15 Jan 2024 0 repositories listed
-
Reliability-Optimized User Admission Control for URLLC Traffic: A Neural Contextual Bandit Approach5 Jan 2024 0 repositories listed
-
Optimal cross-learning for contextual bandits with unknown context distributions3 Jan 2024 0 repositories listed
-
Best-of-Both-Worlds Linear Contextual Bandits27 Dec 2023 0 repositories listed
-
Foundations of Reinforcement Learning and Interactive Decision Making27 Dec 2023 0 repositories listed
-
Diversity-Based Recruitment in Crowdsensing By Combinatorial Multi-Armed Bandits25 Dec 2023 0 repositories listed
-
Zero-Inflated Bandits25 Dec 2023 0 repositories listed
-
Best-of-Both-Worlds Algorithms for Linear Contextual Bandits24 Dec 2023 0 repositories listed
-
Neural Contextual Bandits for Personalized Recommendation21 Dec 2023 0 repositories listed
-
Bayesian Analysis of Combinatorial Gaussian Process Bandits20 Dec 2023 0 repositories listed
-
Distribution-Dependent Rates for Multi-Distribution Learning20 Dec 2023 0 repositories listed
-
Observation-Augmented Contextual Multi-Armed Bandits for Robotic Search and Exploration19 Dec 2023 0 repositories listed
-
Online Restless Multi-Armed Bandits with Long-Term Fairness Constraints16 Dec 2023 0 repositories listed
-
A Hierarchical Nearest Neighbour Approach to Contextual Bandits14 Dec 2023 0 repositories listed
-
Robust and Performance Incentivizing Algorithms for Multi-Armed Bandits with Strategic Agents13 Dec 2023 0 repositories listed
-
Contextual Bandits with Online Neural Regression12 Dec 2023 0 repositories listed
-
Distributed Optimization via Kernelized Multi-armed Bandits7 Dec 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.