ch-universal-bandits

11 Universal Bandit Decision Learning

Partial observation as an information restriction

Bandit feedback changes neither the action object nor the loss class at first. It changes the information channel through which a learner encounters loss. That apparently small change separates three notions that coincide in full information: pathwise recovery, recovery after barycentric averaging, and stability of recovery under infinitesimal variation. This chapter makes those distinctions explicit and then proves when a full-information online guarantee transports through a bandit channel.