lin-0199
Further reading
The reinforcement-learning foundation is Sutton and Barto [ 2018 ] . Temporal difference learning begins with Sutton [ 1988 ] ; gradient-TD methods and their off-policy motivation are developed in Sutton et al. [ 2008 ] , Maei [ 2011 ] . Stochastic Mirror–Prox [ Nemirovski , 2004 , Juditsky et al. , 2011 ] supplies the optimization background for the linear saddle realization.
For decision-relevant credit, compare hindsight credit assignment [ Harutyunyan et al. , 2019 ] , return decomposition [ Arjona-Medina et al. , 2019 ] , and policy-level trust regions [ Schulman et al. , 2015 ] . Natural policy gradient [ Kakade , 2001 ] and natural actor–critic [ Peters and Schaal , 2008 ] provide the information-geometric background for Natural GIRL. Safe policy improvement under limited coverage is a neighboring concern; decision-point restriction provides a recent example [ Sharma et al. , 2025 ] . GIRL differs by tangent-lifting the Bellman factorization and quotienting action-common variation before admission, but the safe and offline RL literature supplies essential baselines for coverage, distribution shift, pessimism, and high-confidence improvement.