ora-0150
12.6.1 Q-learning as asynchronous coordinate computation
12.6.1 Q-learning as asynchronous coordinate computation
The same declaration illuminates Tsitsiklis’s convergence analysis of Q-learning [ Tsitsiklis , 1994 ] . Each state–action pair is a coordinate owner. A visitation event updates one coordinate by a stochastic approximation to the Bellman map,
while other coordinates persist. The index \(\rho (e)\) records the information actually used by the update and may be stale in a parallel realization. Infinite visitation becomes fairness of coordinate ownership; the discounted Bellman contraction supplies the global consistency mechanism; stochastic approximation controls noisy evidence.
This is not a claim that Q-learning itself is universal. It is a structural example showing that an RL convergence proof can factor through an asynchronous information architecture. The new UODL question is which parts of that proof are invariant under changes of schedule presentation and which depend essentially on the scalar norm, contraction modulus, and probability observer.
This chapter’s target results are an assembly theorem for partially nested teams, an obstruction theorem for nonclassical replay, and an observer-valued equilibrium formulation for games. Their infinitesimal forms should recover Theorem 1.30 while exposing the feedback terms that survive in cyclic information structures.