Papers › Multi-objective Good Arm Identification with Bandit Feedback
Multi-objective Good Arm Identification with Bandit Feedback
Xuanke Jiang, Kohei Hatano, Eiji Takimoto
We consider a good arm identification problem in a stochastic bandit setting with multi-objectives, where each arm i∈[K] is associated with M distributions 𝒟ᵢ⁽¹⁾, …, 𝒟ᵢ⁽ᴹ⁾. For each round t, the player/algorithm pulls one arm iₜ and receives a vector feedback, where each component m is sampled according to 𝒟ᵢ⁽ᵐ⁾. The target is twofold, one is finding one arm whose means are larger than the predefined thresholds ξ₁,…,ξ_M with a confidence bound δ and an accuracy rate ϵ with a bounded sample complexity, the other is output $\bot$ to indicate no such arm exists. We propose an algorithm with a sample complexity bound. When M=1 and ϵ= 0, our bound is the same as the one given in the previous work when and novel bounds for M > 1. The proposed algorithm attains better numerical performance than other baselines in the experiments on synthetic and real datasets.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Results from the paper archive 2025-07-28
No leaderboard rows for this paper in the archive.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections