ora-0137

10.7 Quadratic FTRL and typed tangents

Let the mechanism regularizer be

\[ S_\phi (x) = \frac{\lambda }{2}(x-m)^\mathsf TM(x-m), \qquad \lambda {\gt}0,\quad M\succ 0. \]

Assume \(Q_s\succeq 0\). After accumulation, define

\[ H_t=\lambda M+\sum _{s\le t}Q_s, \qquad q_t=\sum _{s\le t}b_s-\lambda Mm. \]

Then \(H_t\succ 0\), and exact-loss FTRL has the unique decision

\begin{equation} x_{t+1}=-H_t^{-1}q_t. \end{equation}
10.7

Proposition 10.6 Typed FTRL action tangents

At the same base decision, evidence and mechanism variations satisfy

\[ H_t\dot x^F_{t+1} =- \left[ \sum _{s\le t}\dot b_s + \left(\sum _{s\le t}\dot Q_s\right)x_{t+1} \right] \]

and

\[ H_t\dot x^\rho _{t+1} =-(\dot q_t^\rho +\dot H_t^\rho x_{t+1}), \]

where

\[ \dot H_t^\rho =\dot\lambda M+\lambda \dot M, \qquad \dot q_t^\rho =-\dot\lambda Mm-\lambda \dot M m-\lambda M\dot m. \]

Their joint first-order response is \(\dot x^F_{t+1}+\dot x^\rho _{t+1}\).

Proof

The first-order condition is

\[ H_tx_{t+1}+q_t=0. \]

Differentiating gives

\[ \dot H_tx_{t+1}+H_t\dot x_{t+1}+\dot q_t=0. \]

Since \(H_t\) is invertible,

\[ \dot x_{t+1} =-H_t^{-1}(\dot q_t+\dot H_tx_{t+1}). \]

Holding the mechanism fixed gives the evidence equation. Holding evidence fixed and differentiating \(\lambda M\) and \(-\lambda Mm\) gives the mechanism equation. Linearity of the differential gives the joint response after the two directions are based at the same decision.

Corollary 10.7 Constant loss shifts are decision-null

Evidence variations of the form

\[ ((0,0,\dot c_s))_{s\le t} \]

lie in the kernel of the FTRL decision tangent.

Proof

The accumulated constant affects only the decision-independent scalar term in the quadratic objective. It is absent from the first-order condition and therefore from Proposition 10.6.

Boundary 10.8 Ambient versus admissible tangents

The tangent calculation occurs in the ambient vector space \(\operatorname {Sym}_d(\mathbb {R})\). If a loss Hessian lies on the boundary of the positive-semidefinite cone, not every ambient direction preserves convexity. An OCO intervention must therefore be restricted to the appropriate tangent cone when convexity preservation is part of the declaration. The readout remains smooth as long as the total \(H_t\) stays positive definite.