ora-0111

8.9 Persistent and online PACC

At time \(t\), let \(\mathcal Q_t\) be the probes licensed by the available information category and let \(P_t\) be their conditional distribution. A persistent learner must control current risk without gratuitously changing answers already settled under earlier doctrines.

Definition 8.9 Persistent PACC

A UOCL execution \((\widehat{\mathcal C}_t)_{t\geq 0}\) is persistently PACC on a settled subdoctrine \(\mathcal Q_0\) when it satisfies the declared PACC risk bounds at each time and its comparison maps \(\widehat{\mathcal C}_t\to \widehat{\mathcal C}_{t+1}\) are conservative on all answers registered as settled in \(\mathcal Q_0\). Under drift, the declaration must additionally specify cumulative or discounted risk and the defect of transporting probes from \(P_t\) to \(P_{t+1}\).

Proposition 8.10 Finite-horizon confidence composition

Suppose that, conditional on every preceding transcript, the learner at each \(t\leq T\) satisfies its PACC bound with failure probability at most \(\delta _t\). Then with probability at least \(1-\sum _{t=1}^T\delta _t\), all \(T\) risk bounds hold simultaneously. If the comparison maps are conservative on the settled doctrine, its registered answers also persist through time \(T\).

Proof

The conditional guarantees imply unconditional failure probabilities bounded by \(\delta _t\). The first claim is the union bound. The second is finite composition of conservative comparison maps, as in Proposition 7.3.

This proposition is modest but clarifies what changes online. A static PACC bound controls a random approximation. Persistent PACC additionally requires typed transport of what has been learned. Adaptive probing calls for concentration under dependent observations; endogenous decisions further require the information and safety semantics developed in Parts IV and VI.