Open AccessMathematicsComputer Science

P. Auer, R. Ortner

2010.10.1Periodica Mathematica Hungarica

DOI: 10.1007/s10998-010-3055-6

tlooto Summary

For this modified UCB algorithm, an improved bound on the regret is given with respect to the optimal reward for K-armed bandits after T trials.

Abstract

Abstract is not available.

Citation format

AUER, P.; ORTNER, R. UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem. Periodica Mathematica Hungarica, 2010, 61: 55–65.