lin-0190

15.3 IL–GTD–MP

In a linear off-policy realization, the infinitesimal-learning variant of gradient temporal difference Mirror–Prox (IL–GTD–MP) augments the GTD–Mirror–Prox objective with a quadratic tangent Bellman penalty [ Sutton et al. , 2008 , Maei , 2011 , Nemirovski , 2004 ] .

Proposition 15.2 Quadratic tangent penalty

For a linear value model and fixed probe covariance, the expected squared tangent Bellman residual is a positive-semidefinite quadratic function of the value parameters. Adding it to the regularized GTD saddle objective preserves monotonicity of the associated variational inequality.

Proof

The Bellman residual derivative is affine in the linear parameters. Its squared norm therefore has a positive-semidefinite Gram Hessian. Adding that term to a monotone regularized GTD operator preserves monotonicity.

Standard stochastic Mirror–Prox rates then apply under the usual boundedness, variance, and step-size hypotheses. The guarantee concerns optimization of the declared linear objective. It does not establish global robustness or policy improvement.

Boundary

A small tangent residual controls first-order variation only on the support and in the metric used to define it. Finite-shift and policy claims additionally need coverage, derivative regularity, Bellman inverse conditioning, and a positive action gap.