Papers › From Theory to Practice with RAVEN-UCB: Addressing Non-Stationarity in Multi-Armed...

From Theory to Practice with RAVEN-UCB: Addressing Non-Stationarity in Multi-Armed Bandits through Variance Adaptation

3 Jun 2025arXiv:2506.02933archive 2025-07-28

Junyi Fang, Yuxun Chen, Yuxin Chen, Chen Zhang

The Multi-Armed Bandit (MAB) problem is challenging in non-stationary environments where reward distributions evolve dynamically. We introduce RAVEN-UCB, a novel algorithm that combines theoretical rigor with practical efficiency via variance-aware adaptation. It achieves tighter regret bounds than UCB1 and UCB-V, with gap-dependent regret of order K σₘₐₓ² logT / Δ and gap-independent regret of order √(K T logT). RAVEN-UCB incorporates three innovations: (1) variance-driven exploration using √(σ̂ₖ² / (Nₖ + 1)) in confidence bounds, (2) adaptive control via αₜ = α₀ / log(t + ϵ), and (3) constant-time recursive updates for efficiency. Experiments across non-stationary patterns - distributional changes, periodic shifts, and temporary fluctuations - in synthetic and logistics scenarios demonstrate its superiority over state-of-the-art baselines, confirming theoretical and practical robustness.

PaperPDFCode

Code

66661654/Raven-UCB officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multi-Armed Bandits

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections