sec-dial-skillopt
4.7 DIAL-SkillOpt: training a discovery skill
The curve-object semantics of Chapter 3 turns the preceding static landscape into a trainable process. SkillOpt treats a skill document as external state for a frozen agent: a separate optimizer uses scored rollouts to propose bounded textual edits, retaining an edit only when it improves held-out validation performance [ Yang et al. , 2026 ] . DIAL-SkillOpt adopts that optimization discipline but changes both the state being maintained and the meaning of success.
A DIAL-SkillOpt discovery skill is a versioned, typed control artifact \(m\) whose compilation determines a state-dependent policy over assimilatory skills, proto-accommodative probes, pairwise-order tests, finite theory- extension proposals, and abstention. Training revises \(m\) using bounded edits and replayable discovery traces. The trained artifact is successful only to the extent that it improves independently certifiable theory construction, not merely the score of its generated prose.
The object being optimized is therefore a policy for discovery. It is not creativity represented as a scalar, nor a differentiable coordinate for the finite sketch map \(J\). The policy learns how to navigate, diagnose, and propose changes to maintained theories; proof, simulation, or new evidence still decides whether any proposed change is admitted.