ora-0116
9.2 Scientific questions for universal online decision learning
UODL is organized by questions that recur whenever categorical identification must support action before all relevant information is available. These questions do not depend on a particular algorithmic family or on a single choice of clock, loss, or comparator.
What makes a decision genuinely online? The learner must act from an information region that is closed under its available past but excludes unrevealed evidence. The first task is therefore to identify the least admissible information object and express non-anticipation as a typed property of the decision map.
When does universal decision semantics persist as evidence grows? Left and right Kan extensions may solve a decision problem at each finite stage without commuting with restriction from one stage to the next. UODL asks for the comparison maps, exactness hypotheses, and invariance conditions under which stagewise solutions form a coherent online object.
How can decisions made with different information be compared? An online learner and a hindsight comparator inhabit different information contexts. A meaningful notion of regret therefore requires a declared comparator object, a transport between information shapes, and an observer that turns their values into an ordered, numerical, or logical defect.
What should happen when an online composition is incomplete or inconsistent? A partial decision history may leave a horn unfilled, admit several coherent fillers, or contain an obstruction to every filler in the current doctrine. The scientific problem is to distinguish ordinary completion, repair within a declaration, and accommodation of the declaration itself.
Which action mechanisms are stable under variation? Feasible decisions may carry convex, metric, order, or richer geometry. Regularized leaders, mirror maps, proximal maps, and related mechanisms become useful categorical constructions only when their existence, sensitivity, and convergence survive the relevant tangent lift.
Which conclusions survive the loss of a global clock or full observation? Bandit feedback, asynchronous computation, and decentralized teams alter the information category rather than merely adding noise to a synchronous protocol. UODL asks which values descend through stochastic channels, which update schedules are equivalent, and which information patterns create irreducible comparison or safety obstructions.
When does learned decision structure remain useful beyond the task that produced it? Persistent structure must transport across changes of environment, observer, and objective while retaining its universal witnesses, repair certificates, and safety conditions. This is the bridge from successful online action to lifelong compositional learning.
Classical online convex optimization supplies a particularly transparent laboratory for these questions: time is a chain, the feasible object is convex, losses are revealed sequentially, and cumulative regret is observed in an ordered additive value space [ Shalev-Shwartz , 2012 , Hazan , 2016 , 2023 ] . Changing any one of these structural choices leads to a different scientific problem rather than merely another entry in a catalog of algorithms. The development below therefore begins with prefix semantics and universal comparison, then studies convex action mechanisms as one exact realization of the broader theory.