ifc-0207
15.13 A grand challenge: constructing a coalgebraic theory of learning
The preceding benchmarks construct operations and presentations inside finite mathematical worlds. A more ambitious test asks whether a learning system can revise the mathematical type of dynamical system through which it understands an environment. Reinforcement learning supplies an unusually clean setting for this challenge. Much of the field is organized around separately named model classes—Markov decision processes, partially observable Markov decision processes, semi-Markov decision processes, and predictive state representations—even though universal coalgebra provides a common language for their unfolding behavior [ Rutten , 2000 ] . Categories for AGI develops this coalgebraic perspective as Universal Reinforcement Learning, or URL [ Mahadevan , 2026b ] . The creative question considered here is not whether an agent can solve one more supplied coalgebra. It is whether the agent can recognize that its current behavior functor is inadequate and construct a better one.
Let \(\mathcal C\) be a category of state objects and let
specify a behavioral type. An environment of that type is an \(F\)-coalgebra
The functor declares what may be observed in one unfolding step: actions, rewards, probability distributions, sets of possible successors, durations, outputs, or combinations of these. A final \(F\)-coalgebra, when available, provides canonical behavior, and coalgebra morphisms and bisimulation compare systems by what they do rather than by the names of their hidden states.
For example, the following equations are schematic finite-state presentations, not assertions that every implementation lives in \(\mathbf{Set}\):
A POMDP adds a hidden-state coalgebra and an observation channel \(X\to \mathcal D O\). A predictive state representation is more naturally read as an observable realization or quotient of unfolding behavior than as merely another hidden-state tuple. These distinctions must be retained: coalgebra organizes behavior, but reward, control, and policy optimization require additional algebraic structure.