ifc-0146

10.7 DIAL-GIRL: the target is an admitted decision state

GIRL supplies tangent diagnostics and quotient-aware learning for sequential decision systems. Its DIAL specialization targets

\[ \mathsf{Target}_{\mathrm{GIRL}} =\bigl((Q,\pi ),\mathbb S_{\mathrm{dec}}, \mathcal C_{\mathrm{decision}}\bigr). \]

The operational object is a critic–policy pair. The declaration records the state and action types, transition support, Bellman and policy-evaluation diagrams, intervention semantics, and quotient geometry. Admission combines structural routing evidence with independent policy return, coherence, and effective-sample-size diagnostics.

In the first realization, \(H_{\mathrm{GIRL}}\) is a within-declaration advantage-quotient critic and actor update. The second operation \(V_{\mathrm{GIRL}}\) transports, forks, restores, or revises the learner state when the regime or its declaration changes. The mixed question is:

\[ \text{transport after learning} \quad \Longleftrightarrow ?\quad \text{learn after transport}. \]

A failure of this comparison does not automatically authorize reset or rollback. It creates a candidate-routing problem.

The DIAL-2 evidence sequence illustrates the distinction. Single-change experiments separated reward and parameter drift from mechanism and topology change. In A–B–A worlds, exact restoration of A was sometimes worse than continual GIRL because B had supplied transferable learning. The paired routing study therefore compared saved, current, and transported states with disjoint proposal and admission trajectories. It selected saved A under antagonistic forgetting and transported learning under positive transfer.

The four implementation levels are consequently distinct. The GIRL tangent and quotient laws present the repair theory; disagreement between learn–then–transport and transport–then–learn supplies the mixed witness. Its integration observer is a finite rollout or recurrence test: does the proposed learner-state transport remain coherent over several Bellman and policy updates, including a return to a previous regime? Only a reproducible failure of this finite test places pressure on the decision sketch itself. Candidate extensions may then introduce a new option, quotient, state abstraction, or reward component and transport prior learner states along that explicit map.

. DIAL-GIRL targets neither maximal return nor recurrence recognition alone. It targets a critic–policy state whose declaration route and deployment value are independently admitted. A remembered state is a proposal, not a rollback command.

The completed DIAL-2 experiments are therefore calibrations of the repair and memory mechanism, not yet demonstrations of creativity. They supply a precursor to RELIC, whose embodied experiments in Chapter 13 move from supplied learner-state alternatives to the construction and admission of temporal decision structure. A stronger DIAL-GIRL experiment would invent a state abstraction, option, skill boundary, reward decomposition, or quotient absent from \(\mathcal R_{\mathrm{GIRL}}\). Reusing a saved policy is adaptation; constructing and admitting a new decision vocabulary is a candidate transformational episode.

A remaining mathematical requirement is quotient-aware functional transport. Raw addition of tabular learner states is coordinate dependent and therefore serves only as a control, not as the intrinsic DIAL-GIRL construction.