ifc-0174

13.2 The RELIC target card

RELIC specializes the DIAL-X contract to sequential decisions. Its three targets are

\[ \mathsf{Target}_{\mathrm{RELIC}} = \bigl(Y_{\mathrm{dec}},\mathbb S_{\mathrm{dec}}, \mathcal C_{\mathrm{dec}}\bigr). \]

The operational target \(Y_{\mathrm{dec}}\) is a deployable policy together with any persistent state or macro-actions it requires. The structural target \(\mathbb S_{\mathrm{dec}}\) declares the types, factorization, temporal interfaces, and observable obligations of that policy. The epistemic target \(\mathcal C_{\mathrm{dec}}\) combines task completion, validity, counter-witnesses, transport, and conservative abstention.

The RELIC creative mandate freezes that decision sketch, its primitive action and observation interfaces, stage grammar, and reference library of admitted policies and schemas. A new composition of registered stages is combinational; a new reachable policy inside the same package is exploratory; and a change to the registered decision package is transformational. Return and novelty remain separate observers.

Operationally, RELIC receives interaction traces, the maintained decision sketch, and the currently registered observation and primitive-action interfaces. It returns a deployable policy together with any persistent state, schema predicate, macro-action, or proposed declaration change that the policy requires. A finite-rollout record and a status-bearing admission record accompany that output; when the evidence cannot distinguish structural inadequacy from ordinary policy error, the correct output is abstention or an evidence request.

In the notation of Chapter 10, \(\mathbb A_{\mathrm{RELIC}}\) presents critic, policy, and action updates inside the maintained decision language, while \(\mathbb B_{\mathrm{RELIC}}\) presents probes and revisions of state, schema, or temporal structure. Their realizations perform within-declaration learning and propose decision-language change, respectively. The observer family records reward, validity, memory, and factorization evidence; \(\Theta _{\mathrm{RELIC}}\) compares the two orders; \(\gamma _{\mathrm{dec}}\) tests finite execution; and \(\mathsf{Adm}_{\mathrm{RELIC}}\) assigns status on disjoint games or traces.

The finite-realization observer \(\gamma _{\mathrm{dec}}\) is distinct from that final certificate family. It asks whether a local policy repair remains coherent over a finite rollout, recurrence, or temporally extended execution, including the state and memory transport required by the candidate decision structure. A one-step reward gain that fails this observer is not eligible for admission.

For the present studies, \(\mathsf{Ctl}_{\mathrm{RELIC}}\) registers current and no-change policies, stronger fixed-declaration learning, generic intrinsic and goal-relevant exploration, isolated structural actions, saved-state restoration, reordered and type-invalid programs, and observation or memory alternatives. Only after the applicable controls fail under the stated contract may RELIC propose a package comparison

\[ \Upsilon _{\mathrm{RELIC}}: \mathfrak R_{\mathrm{RELIC}} \longrightarrow \mathfrak R_{\mathrm{RELIC}}^+. \]

As with skill construction, its presentation component \(J\) does not itself move policy state or memory. A RELIC proposal must also record a supplied or induced semantic transport \(\bar J_{\mathrm{dec}}:M_{\mathrm{dec}}\to M_{\mathrm{dec}}^{+}\), together with the policy-state and memory lift over it.

The RELIC artifact dossier consequently contains more than the deployed policy. It records the constructed temporal composition or task-stage operator, its package comparison \(\Upsilon _{\mathrm{RELIC}}\), the state and memory transported into the new policy language, disjoint-game admission and confirmation results, and a replayable version of the resulting program. The strongest experiments below admit such constructions up to behavioral equivalence inside registered primitive actions and stage contracts; they do not establish invention of a new primitive action or observation type.

For the experiments in this chapter, the horizontal direction is learning or acting inside the current policy language. The vertical, proto-accommodative direction probes or proposes a change to the language: a state coordinate, schema predicate, or temporally extended composition. Only a finite proposal that passes realization and admission is installed as accommodation. Their mixed question is operational:

\[ \begin{aligned} & \text{learn and act in the old declaration, then extend it}\\ & \hspace{4em}\Longleftrightarrow ?\quad \text{extend the declaration, then learn and act}. \end{aligned} \]

A repeated action, Bellman inconsistency, non-Markov residual, or failed factorization can make the disagreement observable. None of these witnesses automatically licenses accommodation. It localizes pressure on the current model and triggers a proposal, finite-realization, and admission process.

The RELIC loop. Creativity is not identified with the obstruction or the proposal. A candidate must first survive a registered finite rollout or recurrence test, and then an independent admission gate, before it becomes persistent policy structure.
Figure 13.1 The RELIC loop. Creativity is not identified with the obstruction or the proposal. A candidate must first survive a registered finite rollout or recurrence test, and then an independent admission gate, before it becomes persistent policy structure.