Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Tight Regret Bounds for Stochastic Combinatorial Semi-Bandits

Jan 27, 2015

Branislav Kveton, Zheng Wen, Azin Ashkan, Csaba Szepesvari

Figure 1 for Tight Regret Bounds for Stochastic Combinatorial Semi-Bandits

Share this with someone who'll enjoy it:

Abstract:A stochastic combinatorial semi-bandit is an online learning problem where at each step a learning agent chooses a subset of ground items subject to constraints, and then observes stochastic weights of these items and receives their sum as a payoff. In this paper, we close the problem of computationally and sample efficient learning in stochastic combinatorial semi-bandits. In particular, we analyze a UCB-like algorithm for solving the problem, which is known to be computationally efficient; and prove $O(K L (1 / \Delta) \log n)$ and $O(\sqrt{K L n \log n})$ upper bounds on its $n$-step regret, where $L$ is the number of ground items, $K$ is the maximum number of chosen items, and $\Delta$ is the gap between the expected returns of the optimal and best suboptimal solutions. The gap-dependent bound is tight up to a constant factor and the gap-free bound is tight up to a polylogarithmic factor.

* Proceedings of the 18th International Conference on Artificial Intelligence and Statistics

View paper on

Share this with someone who'll enjoy it:

Title:Tight Regret Bounds for Stochastic Combinatorial Semi-Bandits

Paper and Code