ora-0140
11.1 The bandit information channel
Let \(A\) be a finite arm object with \(K=|A|\), let \(\Delta _A=D(A)\) be its free convex algebra of randomized actions, and put \(\mathcal L=[0,1]^A\). For \(p\in \Delta _A\) and \(\ell \in \mathcal L\), the stochastic observation channel is the Kleisli arrow
The registered evidence is therefore the sampled arm and its scalar loss, not the complete loss vector.
A finite-arm bandit UODL declaration consists of:
a history filtration \((\mathcal F_t)_{t\geq 0}\);
an \(\mathcal F_{t-1}\)-measurable, full-support action distribution \(p_t\in \Delta _A^\circ \);
a loss \(\ell _t\in \mathcal L\) fixed before the current sample \(A_t\sim p_t\);
the channel \(\kappa _{p_t}\), together with a declared evidence repair; and
a comparator sketch and an observer that values the comparison.
An oblivious declaration fixes the loss sequence in advance. A predictable adaptive declaration permits \(\ell _t\) to depend on \(\mathcal F_{t-1}\). An adversary that observes \(A_t\) before choosing \(\ell _t\) lies outside this declaration.
For \(p\in \Delta _A^\circ \), define the importance repair
For every full-support \(p\in \Delta _A^\circ \), importance repair is a left inverse to the bandit observation channel after applying the convex-algebra expectation:
Equivalently, \(\mathbb E_{a\sim p}[R_p(a,\ell (a))]=\ell \).
For each coordinate \(b\in A\), barycentric evaluation gives
The equality holds coordinatewise in the product convex algebra \(\mathcal L\).