ora-0140

11.1 The bandit information channel

Let \(A\) be a finite arm object with \(K=|A|\), let \(\Delta _A=D(A)\) be its free convex algebra of randomized actions, and put \(\mathcal L=[0,1]^A\). For \(p\in \Delta _A\) and \(\ell \in \mathcal L\), the stochastic observation channel is the Kleisli arrow

\begin{equation} \kappa _p(\ell ) =\sum _{a\in A}p(a)\, \delta _{(a,\ell (a))} \in D(A\times [0,1]). \end{equation}
11.1

The registered evidence is therefore the sampled arm and its scalar loss, not the complete loss vector.

Definition 11.1 Finite-arm bandit UODL declaration

A finite-arm bandit UODL declaration consists of:

  1. a history filtration \((\mathcal F_t)_{t\geq 0}\);

  2. an \(\mathcal F_{t-1}\)-measurable, full-support action distribution \(p_t\in \Delta _A^\circ \);

  3. a loss \(\ell _t\in \mathcal L\) fixed before the current sample \(A_t\sim p_t\);

  4. the channel \(\kappa _{p_t}\), together with a declared evidence repair; and

  5. a comparator sketch and an observer that values the comparison.

An oblivious declaration fixes the loss sequence in advance. A predictable adaptive declaration permits \(\ell _t\) to depend on \(\mathcal F_{t-1}\). An adversary that observes \(A_t\) before choosing \(\ell _t\) lies outside this declaration.

For \(p\in \Delta _A^\circ \), define the importance repair

\begin{equation} R_p(a,y)(b)=\frac{y\, \mathbf1\{ a=b\} }{p(a)}, \qquad b\in A. \end{equation}
11.2

Bandit evidence is not inverted sample by sample. The channel is split only after lifting repaired observations into the distribution monad and applying its barycentric algebra.
Figure 11.1 Bandit evidence is not inverted sample by sample. The channel is split only after lifting repaired observations into the distribution monad and applying its barycentric algebra.
Proposition 11.2 Barycentric reconstruction of bandit evidence

For every full-support \(p\in \Delta _A^\circ \), importance repair is a left inverse to the bandit observation channel after applying the convex-algebra expectation:

\[ \beta _{\mathcal L}\circ D(R_p)\circ \kappa _p =1_{\mathcal L}. \]

Equivalently, \(\mathbb E_{a\sim p}[R_p(a,\ell (a))]=\ell \).

Proof

For each coordinate \(b\in A\), barycentric evaluation gives

\[ \sum _{a\in A}p(a) \frac{\ell (a)\mathbf1\{ a=b\} }{p(a)}=\ell (b). \]

The equality holds coordinatewise in the product convex algebra \(\mathcal L\).