ora-0143
11.4 The expectation-level regret-transfer theorem
The previous results determine the observer through which a full-information theorem may be transported. Let \(\widehat\ell _t=R_{p_t}(A_t,\ell _t(A_t))\).
Let \(\ell _1,\ldots ,\ell _T\in [0,1]^A\) be an oblivious loss sequence, and let \(p_t\in \Delta _A^\circ \) be adapted to \(\mathcal F_{t-1}\). Suppose a full-information online rule applied to \(\widehat\ell _t\) satisfies, on every realized surrogate sequence and for every fixed \(u\) in a comparator class \(\mathcal C\subseteq \Delta _A\),
Then its bandit realization satisfies
Pathwise, \(\langle p_t,\widehat\ell _t\rangle =\ell _t(A_t)\). By adaptedness and barycentric reconstruction,
for every fixed \(u\). Take expectations in the assumed pathwise inequality and then the infimum over \(u\in \mathcal C\).
For a predictable adaptive adversary, the same calculation yields pseudo-regret against each comparator fixed independently of the realized loss sequence. It does not by itself justify moving an expectation through a random hindsight infimum. If the adversary sees the current sampled arm before fixing the loss, even conditional barycentric reconstruction fails. These are changes of declaration, not technical footnotes.