ch-girl
15 Infinitesimal Reinforcement Learning
GIRL
Temporal-difference learning turns Bellman inconsistency into a scalar error. GIRL retains the diagram that produced that error. It distinguishes the base Bellman obstruction, its tangent sensitivity, the action-common variation that should be quotiented away, the computational branch carrying the failure, and the evidence required to transport a repair across regimes.
11. GIRL abbreviates Gradient Infinitesimal Reinforcement Learning. It is an architectural layer above TD, not a replacement for the underlying policy-evaluation algorithm. ↩