ch-girl

15 Infinitesimal Reinforcement Learning

GIRL

Temporal-difference learning turns Bellman inconsistency into a scalar error. GIRL retains the diagram that produced that error. It distinguishes the base Bellman obstruction, its tangent sensitivity, the action-common variation that should be quotiented away, the computational branch carrying the failure, and the evidence required to transport a repair across regimes.

11. GIRL abbreviates Gradient Infinitesimal Reinforcement Learning. It is an architectural layer above TD, not a replacement for the underlying policy-evaluation algorithm.