Open AccessMathematicsComputer Science
P. Auer, R. Ortner
2010.10.1Periodica Mathematica Hungarica
tlooto Summary
For this modified UCB algorithm, an improved bound on the regret is given with respect to the optimal reward for K-armed bandits after T trials.
Abstract
Abstract is not available.
Citation format
AUER, P.; ORTNER, R. UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem. Periodica Mathematica Hungarica, 2010, 61: 55–65.