ora-0143

11.4 The expectation-level regret-transfer theorem

The previous results determine the observer through which a full-information theorem may be transported. Let \(\widehat\ell _t=R_{p_t}(A_t,\ell _t(A_t))\).

Theorem 11.6 Bandit regret transfer through barycentric repair

Let \(\ell _1,\ldots ,\ell _T\in [0,1]^A\) be an oblivious loss sequence, and let \(p_t\in \Delta _A^\circ \) be adapted to \(\mathcal F_{t-1}\). Suppose a full-information online rule applied to \(\widehat\ell _t\) satisfies, on every realized surrogate sequence and for every fixed \(u\) in a comparator class \(\mathcal C\subseteq \Delta _A\),

\[ \sum _{t=1}^T\langle p_t,\widehat\ell _t\rangle -\sum _{t=1}^T\langle u,\widehat\ell _t\rangle \leq B_T. \]

Then its bandit realization satisfies

\[ \mathbb E\! \left[\sum _{t=1}^T\ell _t(A_t)\right] -\inf _{u\in \mathcal C}\sum _{t=1}^T\langle u,\ell _t\rangle \leq B_T. \]
Proof

Pathwise, \(\langle p_t,\widehat\ell _t\rangle =\ell _t(A_t)\). By adaptedness and barycentric reconstruction,

\[ \mathbb E[\widehat\ell _t\mid \mathcal F_{t-1}]=\ell _t, \qquad \mathbb E\langle u,\widehat\ell _t\rangle =\langle u,\ell _t\rangle \]

for every fixed \(u\). Take expectations in the assumed pathwise inequality and then the infimum over \(u\in \mathcal C\).

For a predictable adaptive adversary, the same calculation yields pseudo-regret against each comparator fixed independently of the realized loss sequence. It does not by itself justify moving an expectation through a random hindsight infimum. If the adversary sees the current sampled arm before fixing the loss, even conditional barycentric reconstruction fails. These are changes of declaration, not technical footnotes.