ora-0142
11.3 Exploration as infinitesimal admissibility
For \(\alpha {\gt}0\), write
This is not merely an algorithmic restriction. It is a domain on which the repair map admits controlled tangent variation.
Along a differentiable curve \(\varepsilon \mapsto (p_\varepsilon ,y_\varepsilon )\) with \(p_0=p\in \Delta _A^\alpha \), the conditional repaired evidence obeys
For \(0\leq y\leq 1\),
If \(\widehat\ell =R_p(A,\ell (A))\) with \(A\sim p\), then
Consequently, mixing any policy \(q\) with a full-support reference law
is a tangent-domain repair. If \(\alpha =\gamma \min _a u(a)\), then the action mechanism lies in \(\Delta _A^\alpha \). The usual exploration–variance tradeoff therefore has a LINCS interpretation: exploration controls whether the evidence reconstruction is infinitesimally admissible.
Let \(p_\varepsilon \in \Delta _A^\circ \) and \(\ell _\varepsilon \in \mathcal L\) be differentiable curves. Then
Thus variations of the sampling mechanism and of the loss evidence cancel correctly after barycentric reconstruction, even though neither is stable pathwise near the simplex boundary.
The composite inside the derivative equals \(\ell _\varepsilon \) for every \(\varepsilon \) by 11.2. Differentiating this identity proves the claim. Coordinatewise, the derivative of the sampling weight cancels the derivative of its reciprocal in the importance repair.