ora-0133

10.5 The FTRL action normal form

At time \(t\), an OCO learner selects \(x_t\in K\), observes a convex loss \(\ell _t\), and incurs \(\ell _t(x_t)\). Its static regret against a fixed comparator is

\[ R_T = \sum _{t=1}^T\ell _t(x_t) - \inf _{u\in K}\sum _{t=1}^T\ell _t(u). \]

Our first theorem concerns accumulation and action selection, not the right-stage semantics of more general comparator classes. The constant comparator and numerical difference factor through the observer declaration of 9.13.

A broad FTRL template is

\begin{equation} x_{t+1} = \operatorname *{arg\, min}_{x\in K} \left\{ \langle g_{1:t},x\rangle +A_t^\Psi (x) +S_t(x) \right\} , \end{equation}
10.2

where \(g_{1:t}\) is accumulated smooth-loss evidence, \(A_t^\Psi \) declares the treatment of a nonsmooth composite term, and \(S_t\) declares stabilizing geometry. The minimization is the action readout.

McMahan’s comparison [ McMahan , 2011 ] exposes two independent design axes:

Treatment of \(\Psi \)

Origin-centered stabilizer

Proximal stabilizer

Cumulative exact

RDA

FTRL-Proximal

Earlier terms linearized

AOGD

FOBOS

This grid will later support controlled LINCS interventions. Here we first calibrate the smooth quadratic interior.