sec-observer-regret
9.13 Observers and the origin of regret
An online decision and a hindsight comparator inhabit different information semantics. The online branch selects \(x_t\) from the principal past \(\downarrow t\). The comparator branch sees the complete horizon but is restricted to a declared class, such as one action held fixed at every time. Neither branch alone is a regret. A comparison exists only after an observer places their evaluated outcomes in a common value type.
Let \(\mathbb T_T=\{ 1{\lt}\cdots {\lt}T\} \), let \(X\) be an action object, and let \(\pi :\mathbb T_T\to 1\) be the terminal functor. Precomposition gives the constant-diagram functor
with right adjoint
Assume \(\mathbb T_T\) is nonempty and connected and the required limits exist. A trajectory object \(c\in [\mathbb T_T,\mathcal D]\) is isomorphic to a constant section precisely when the counit
is an isomorphism. Thus the static comparator declaration is the full subcategory on the right-Kan-consistent trajectories.
If \(c\cong \Delta u\), connectedness gives \(\operatorname {Ran}_\pi c\cong u\), and the counit is the constant-diagram isomorphism. Conversely, an invertible counit exhibits \(c\cong \Delta \operatorname {Ran}_\pi c\), so \(c\) is isomorphic to a constant section.
For a fixed action carrier \(X\), a comparator trajectory is a natural section \(\Delta 1\to \Delta X\). Connectedness forces all of its temporal components to be the same generalized element \(u:1\to X\). The right Kan condition therefore declares staticness; minimizing cumulative loss over those sections is a subsequent decision readout.
Now let \((V,\oplus ,0,\leq )\) be an ordered commutative value monoid, let \(e_t\) evaluate the revealed information and an action in \(V\), and let \(\sigma _T:V^T\to V\) be the declared accumulation map. The two branches produce
The first line is a UODL quantity because each \(x_t\) is adapted to \(F_t\). The second is a UDL quantity computed from the full diagram \(F_T\), with right-Kan consistency restricting the comparator trajectory to a constant section. It is a hindsight semantic object, not an executable online policy.
An online comparison observer is a declared morphism
whose two inputs are the cumulative values of an adapted UODL branch and a full-information UDL reference branch. Its output is the observable comparison object.
When \(V\) is an ordered abelian group and \(W=V\), the numerical regret observer is
If \(V\) is only ordered, the observer may return the proposition \(b\leq a\). In a residuated value object it may instead return a residual. Thus subtraction is not supplied by UDL or UODL themselves; it is extra observer structure.
This factorization also separates two uses of “second term.” In the regret difference, \(L_T^{\mathrm{stat}}\) is the full-information UDL term. In the FTRL action objective, by contrast, the regularizer is a mechanism term. It stabilizes the adapted action readout and is not the hindsight comparator. Standard static regret therefore has the typed architecture