ch-online-persistent-udl
9 Online and Persistent Universal Decision Learning
From UOCL specialization to prefix semantics, nerves, homotopy coherence, and repair
ORACLE states the semantic problem of identifying a predictive categorical world, and UOCL supplies its effective online machinery. Universal Online Decision Learning begins when a UOCL hypothesis is equipped with information, action, consistency, and observation maps. Online convex optimization (OCO) then describes one sequential decision problem over a convex feasible set under convex losses revealed after each action [ Hazan , 2023 ] . Universal Online Decision Learning asks a structural question:
Which parts of an online convex decision problem are determined by universal properties, and when do those constructions remain stable under infinitesimal variation?
The operadic theory of convexity characterizes convex sets as algebras over a convexity PROP and equips their category with additional monoidal structure [ Haderi et al. , 2024 ] . UODL uses this object-level account for the feasible decision domain. It then adds Universal Decision Learning (UDL) for extension and consistency, and the LINCS discipline for typed infinitesimal comparison and repair [ Mahadevan , 2021 , 2026a , 2026c ] .
Thus UODL is not the ambient theory of this book; it is the decision-bearing specialization of UOCL, restricted to categories carrying an online decision declaration. Its first technical problem should be simple enough for exact calculation but rich enough to distinguish evidence, action geometry, and consistency. Online convex optimization (OCO) supplies this setting [ Shalev-Shwartz , 2012 , Hazan , 2016 ] . Many OCO algorithms admit FTRL interpretations, and mirror-descent and FTRL presentations can be exactly equivalent under explicit hypotheses [ McMahan , 2011 ] . This makes FTRL a candidate normal form for the action mechanism, not an assertion that all optimization algorithms are secretly identical.
Contributions.
The UODL specialization currently establishes the following contributions.
It defines online and persistent UDL using principal-past restriction, non-anticipation, and coherent semantic transport, and proves a fixed-shape lifting theorem for Kan invariance.
It factors regret through a causal UODL branch, a full-information UDL comparator branch, and a separately typed observer, while identifying static trajectories through a right-Kan counit condition. It then defines information-relative intrinsic regret, recovers ordinary OCO on a classical chain, identifies the replay obstruction for nonclassical information, and realizes crossword and Sudoku solving as decentralized limit problems. It then compares the universal Sudoku solution object with the learned fixed-point dynamics of a Flow Reasoning Model.
It defines comparator sketches as right-Kan descent data, proves an essential-image representation theorem, constructs the static-to-dynamic refinement spectrum, and separates universal descent defects from their metric and numerical observation.
It represents finite decision histories by a simplicial nerve, distinguishes external time from compositional depth, and gives a sufficient condition for online-to-offline UDL continuity.
It defines homotopy-coherent online UDL after localization at semantic weak equivalences, interprets horn filling as a space of repairs, gives a two-component inner-horn example in which later evidence forces a repair, and proves derived skeletal continuity for finite consistency shapes in a stable target.
It defines fixed-stratum infinitesimal intrinsic models and proves that acyclic information dependence makes the linearized closed-loop operator nilpotent, yielding a unique tangent response by causal forward substitution.
It gives a categorical OCO declaration in which the feasible domain is a convex algebra and convex losses are order-lax maps, and proves that its finite-dimensional Euclidean realization recovers the standard OCO protocol.
It separates FTRL into accumulated evidence, composite-loss treatment, stabilizing geometry, and an optimization readout.
It proves a finite enriched left-Kan accumulation theorem and its exact tangent lift, then derives typed FTRL decision sensitivities and a decision-null quotient.
It separates base convergence, tangent stability, and preservation of asymptotic limits, proving sufficient lifting results for normalized quadratic FTRL and contractive smooth update systems.
It presents FTRL as a metric proximal map, proves its smooth tangent lift and contraction bound, and identifies the nonsmooth boundary where graphical rather than ordinary tangents are required.
It models bandit feedback as a stochastic information channel, separates pathwise impossibility from barycentric reconstruction, proves tangent admissibility and regret transfer, recovers the adversarial finite-arm rate, and identifies smoothed tangent covectors as the repair target in bandit convex optimization.
It separates an intrinsic asynchronous event category from its external serializations, proves schedule invariance when incomparable updates commute, and carries the example through distributed minimization, asynchronous Q-learning, and stale-authority safety.
It identifies free convex completion as the universal carrier underlying categorical online boosting.
It formulates agentic safety as a property of the realized information structure, proves compositional safety from generator-level admission and safety transport under pointwise right-Kan extension, and identifies faithful audit as a necessary condition for observer-based enforcement. It then interprets temporal behavior types as trust contracts and proves local-to-global verification by sheaf descent.
The scope is intentionally narrow. We do not yet identify comparator regret with a particular right Kan extension, derive a new regret-optimal algorithm, or claim that a universal tangent comparison is causal.