Browse State-of-the-Art › Multi-Armed Bandits › Papers, page 4
Multi-Armed Bandits
Papers archive 2025-07-28
archive papers tagged: 1,262 · with a code link: 253 · where Syntology ran a sample: 55 (44 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (55 of 1,262 tagged: 44 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 4 of 13: papers 301 to 400 of 1,262, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Towards Understanding the Benefit of Multitask Representation Learning in Decision Process1 Mar 2025 0 repositories listed
-
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models27 Feb 2025 0 repositories listed
-
Heterogeneous Multi-Agent Bandits with Parsimonious Hints22 Feb 2025 0 repositories listed
-
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling20 Feb 2025 0 repositories listed
-
Continuous K-Max Bandits19 Feb 2025 0 repositories listed
-
Efficient and Optimal Policy Gradient Algorithm for Corrupted Multi-armed Bandits19 Feb 2025 0 repositories listed
-
Contextual Linear Bandits with Delay as Payoff18 Feb 2025 0 repositories listed
-
Model selection for behavioral learning data and applications to contextual bandits18 Feb 2025 0 repositories listed
-
Near-Optimal Private Learning in Linear Contextual Bandits18 Feb 2025 0 repositories listed
-
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing15 Feb 2025 0 repositories listed
-
Heterogeneous Multi-agent Multi-armed Bandits on Stochastic Block Models11 Feb 2025 0 repositories listed
-
Provably Efficient RLHF Pipeline: A Unified View from Contextual Bandits11 Feb 2025 0 repositories listed
-
Quantile Multi-Armed Bandits with 1-bit Feedback10 Feb 2025 0 repositories listed
-
Towards a Sharp Analysis of Offline Policy Learning for f-Divergence-Regularized Contextual Bandits9 Feb 2025 0 repositories listed
-
Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback7 Feb 2025 0 repositories listed
-
Early Stopping in Contextual Bandits and Inferences5 Feb 2025 0 repositories listed
-
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards4 Feb 2025 0 repositories listed
-
Nearly Tight Bounds for Exploration in Streaming Multi-armed Bandits with Known Optimality Gap3 Feb 2025 0 repositories listed
-
Optimizing Online Advertising with Multi-Armed Bandits: Mitigating the Cold Start Problem under Auction Dynamics3 Feb 2025 0 repositories listed
-
Meta-Prompt Optimization for LLM-Based Sequential Decision Making2 Feb 2025 0 repositories listed
-
Multi-agent Multi-armed Bandit with Fully Heavy-tailed Dynamics31 Jan 2025 0 repositories listed
-
Nearly-Optimal Bandit Learning in Stackelberg Games with Side Information31 Jan 2025 0 repositories listed
-
Offline Learning for Combinatorial Multi-armed Bandits31 Jan 2025 0 repositories listed
-
Contextual Online Decision Making with Infinite-Dimensional Functional Regression30 Jan 2025 0 repositories listed
-
Breaking the log(1/Δ₂) Barrier: Better Batched Best Arm Identification with Adaptive Grids29 Jan 2025 0 repositories listed
-
HD-CB: The First Exploration of Hyperdimensional Computing for Contextual Bandits Problems28 Jan 2025 0 repositories listed
-
Restless Multi-armed Bandits under Frequency and Window Constraints for Public Service Inspections27 Jan 2025 0 repositories listed
-
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy24 Jan 2025 0 repositories listed
-
Optimal Multi-Objective Best Arm Identification with Fixed Confidence23 Jan 2025 0 repositories listed
-
Efficient Implementation of LinearUCB through Algorithmic Improvements and Vector Computing Acceleration for Embedded Learning Systems22 Jan 2025 0 repositories listed
-
Heterogeneous Multi-Player Multi-Armed Bandits Robust To Adversarial Attacks21 Jan 2025 0 repositories listed
-
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness20 Jan 2025 0 repositories listed
-
Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy17 Jan 2025 0 repositories listed
-
Neural Risk-sensitive Satisficing in Contextual Bandits15 Jan 2025 0 repositories listed
-
Differentially Private Kernelized Contextual Bandits13 Jan 2025 0 repositories listed
-
Finite-Horizon Single-Pull Restless Bandits: An Efficient Index Policy For Scarce Resource Allocation10 Jan 2025 0 repositories listed
-
On The Statistical Complexity of Offline Decision-Making10 Jan 2025 0 repositories listed
-
An Instrumental Value for Data Production and its Application to Data Pricing24 Dec 2024 0 repositories listed
-
A Novel Approach to Balance Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes and its Implementation in BEACON23 Dec 2024 0 repositories listed
-
Lagrangian Index Policy for Restless Bandits with Average Reward17 Dec 2024 0 repositories listed
-
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization16 Dec 2024 0 repositories listed
-
An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints11 Dec 2024 0 repositories listed
-
Conservative Contextual Bandits: Beyond Linear Representations9 Dec 2024 0 repositories listed
-
UCB algorithms for multi-armed bandits: Precise regret and adaptive inference9 Dec 2024 0 repositories listed
-
Coordinated Multi-Armed Bandits for Improved Spatial Reuse in Wi-Fi4 Dec 2024 0 repositories listed
-
Data Acquisition for Improving Model Fairness using Reinforcement Learning4 Dec 2024 0 repositories listed
-
Selective Reviews of Bandit Problems in AI via a Statistical View3 Dec 2024 0 repositories listed
-
Contextual Bandits in Payment Processing: Non-uniform Exploration and Supervised Learning at Adyen30 Nov 2024 0 repositories listed
-
Achieving PAC Guarantees in Mechanism Design through Multi-Armed Bandits30 Nov 2024 0 repositories listed
-
Off-policy estimation with adaptively collected data: the power of online learning19 Nov 2024 0 repositories listed
-
Multi-Agent Stochastic Bandits Robust to Adversarial Corruptions12 Nov 2024 0 repositories listed
-
Individual Regret in Cooperative Stochastic Multi-Armed Bandits10 Nov 2024 0 repositories listed
-
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF7 Nov 2024 0 repositories listed
-
Structure Matters: Dynamic Policy Gradient7 Nov 2024 0 repositories listed
-
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset6 Nov 2024 0 repositories listed
-
Rising Rested Bandits: Lower Bounds and Efficient Algorithms6 Nov 2024 0 repositories listed
-
MBExplainer: Multilevel bandit-based explanations for downstream models with augmented graph embeddings1 Nov 2024 0 repositories listed
-
FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation26 Oct 2024 0 repositories listed
-
Learning to Explore with Lagrangians for Bandits under Unknown Linear Constraints24 Oct 2024 0 repositories listed
-
Optimal Streaming Algorithms for Multi-Armed Bandits23 Oct 2024 0 repositories listed
-
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits21 Oct 2024 0 repositories listed
-
Contextual Bandits with Arm Request Costs and Delays17 Oct 2024 0 repositories listed
-
Is Prior-Free Black-Box Non-Stationary Reinforcement Learning Feasible?17 Oct 2024 0 repositories listed
-
How Does Variance Shape the Regret in Contextual Bandits?16 Oct 2024 0 repositories listed
-
Comparative Performance of Collaborative Bandit Algorithms: Effect of Sparsity and Exploration Intensity15 Oct 2024 0 repositories listed
-
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing14 Oct 2024 0 repositories listed
-
Contextual Bandits with Non-Stationary Correlated Rewards for User Association in MmWave Vehicular Networks8 Oct 2024 0 repositories listed
-
Diminishing Exploration: A Minimalist Approach to Piecewise Stationary Multi-Armed Bandits8 Oct 2024 0 repositories listed
-
Stochastic Bandits for Egalitarian Assignment8 Oct 2024 0 repositories listed
-
7 Oct 2024 0 repositories listed Syntology 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
High Probability Bound for Cross-Learning Contextual Bandits with Unknown Context Distributions5 Oct 2024 0 repositories listed
-
Minimax-optimal trust-aware multi-armed bandits4 Oct 2024 0 repositories listed
-
Online Posterior Sampling with a Diffusion Prior4 Oct 2024 0 repositories listed
-
uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs4 Oct 2024 0 repositories listed
-
On Lai's Upper Confidence Bound in Multi-Armed Bandits3 Oct 2024 0 repositories listed
-
Fast and Sample Efficient Multi-Task Representation Learning in Stochastic Contextual Bandits2 Oct 2024 0 repositories listed
-
Stabilizing the Kumaraswamy Distribution1 Oct 2024 0 repositories listed
-
Optimism in the Face of Ambiguity Principle for Multi-Armed Bandits30 Sep 2024 0 repositories listed
-
Linear Contextual Bandits with Interference24 Sep 2024 0 repositories listed
-
Second Order Bounds for Contextual Bandits with Function Approximation24 Sep 2024 0 repositories listed
-
Designing an Interpretable Interface for Contextual Bandits23 Sep 2024 0 repositories listed
-
Causal Feature Selection Method for Contextual Multi-Armed Bandits in Recommender System20 Sep 2024 0 repositories listed
-
Partially Observable Contextual Bandits with Linear Payoffs17 Sep 2024 0 repositories listed
-
A Hybrid Meta-Learning and Multi-Armed Bandit Approach for Context-Specific Multi-Objective Recommendation Optimization13 Sep 2024 0 repositories listed
-
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits13 Sep 2024 0 repositories listed
-
Batched Online Contextual Sparse Bandits with Sequential Inclusion of Features13 Sep 2024 0 repositories listed
-
Modified Meta-Thompson Sampling for Linear Bandits and Its Bayes Regret Analysis10 Sep 2024 0 repositories listed
-
Faster Q-Learning Algorithms for Restless Bandits6 Sep 2024 0 repositories listed
-
Whittle Index Learning Algorithms for Restless Bandits with Constant Stepsizes6 Sep 2024 0 repositories listed
-
Improving Thompson Sampling via Information Relaxation for Budgeted Multi-armed Bandits28 Aug 2024 0 repositories listed
-
Contextual Bandit with Herding Effects: Algorithms and Recommendation Applications26 Aug 2024 0 repositories listed
-
Representative Arm Identification: A fixed confidence approach to identify cluster representatives26 Aug 2024 0 repositories listed
-
Online Fair Division with Contextual Bandits23 Aug 2024 0 repositories listed
-
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards22 Aug 2024 0 repositories listed
-
Dynamic Product Image Generation and Recommendation at Scale for Personalized E-commerce22 Aug 2024 0 repositories listed
-
Multi-agent Multi-armed Bandits with Stochastic Sharable Arm Capacities20 Aug 2024 0 repositories listed
-
Contextual Bandits for Unbounded Context Distributions19 Aug 2024 0 repositories listed
-
GINO-Q: Learning an Asymptotically Optimal Index Policy for Restless Multi-armed Bandits19 Aug 2024 0 repositories listed
-
Reciprocal Learning12 Aug 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.