ora-0146
Further reading
The adversarial finite-arm construction and its classical regret analysis originate with Auer et al. [ 2002 ] . The one-point smoothed-gradient method for bandit convex optimization is due to Flaxman et al. [ 2005 ] . For a comprehensive modern treatment spanning stochastic, adversarial, Bayesian, contextual, linear, and combinatorial bandits, see Lattimore and Szepesvári [ 2020 ] . Hazan [ 2023 ] develops both finite-arm and convex bandit optimization within the broader OCO landscape. The information-channel, barycentric splitting, tangent-admissibility, and observer distinctions in this chapter are the UODL reformulation.