ifc-0182

13.12 Constructing a missing stage

The three studies in this section progressively refine the observer rather than merely enlarge the search. B5 constructs programs from registered action fragments but cannot separate its selected ordering from a reordered control. B5.1 actively acquires disagreement states, yet its raw syntax space contains programs that remain behaviorally indistinguishable. B5.2 quotients that space by observed action traces and ranks distinct classes with typed coherence before reward and efficiency. The first two studies abstain; the third admits a constructed stage up to the resulting empirical equivalence relation.

Stage construction removes a stronger piece of scaffolding. B4 ordered known stages; B5 deliberately removes \(\mathsf{LocateTarget}\) from the registered task-stage vocabulary. No permutation of the remaining operators can satisfy the acquisition precondition coherently. Because that operator inventory is finite and exhaustively typed, this impossibility is the registered nonrealizability gate for the rung; empirical search failure alone would not have been sufficient. The proposal mechanism instead receives a finite grammar of lower-level action fragments and must construct a new operator of type

\[ \mathsf{SearchState}[\mathsf{target\ not\ actionable}] \longrightarrow \mathsf{SearchState}[\mathsf{target\ actionable}]. \]

Its observable codomain witness is an admissible take command naming the target. The grammar contains opening an unvisited receptacle, navigating to an unvisited location, navigating away from the destination, and navigating to the destination. Twenty-four three-fragment priority programs are evaluated on five proposal games. The selected program is frozen before eight disjoint admission games are opened.

The proposal phase selected

\[ \mathsf{Priority}(\mathsf{openUnvisited}, \mathsf{goUnvisitedAny},\mathsf{goNonDestination}), \]

with 5/5 task success and 5/5 contract witnesses. On admission it solved 6/8 games, compared with 4/8 for the missing-stage control and 0/8 for the frozen DQN. It satisfied its codomain witness on 7/8 games and incurred neither invalid actions nor precondition violations. A control with an incorrect stop predicate solved 4/8 but accumulated 107 precondition violations. These comparisons show why task reward alone cannot certify a declaration extension.

B5 nevertheless abstained. Reversing the selected fragment order still solved 5/8 games, leaving a gain of only \(1/8\), below the preregistered \(2/8\) margin. The untouched confirmation partition was therefore not opened. The result supports the usefulness of adding a typed target-search stage, but does not identify the internal priority structure sharply enough for admission. Several programs tied on proposal success and contract satisfaction; the available worlds did not provide counter-witnesses that separated them.

Experiment: RELIC–ALFWorld-B5. Expressive obstruction: \(\mathsf{LocateTarget}\) withheld from the task-stage vocabulary.
Construction language: 24 priority programs built from four registered navigation and inspection fragments.
Frozen candidate: \(\mathsf{openUnvisited}\), then \(\mathsf{goUnvisitedAny}\), then \(\mathsf{goNonDestination}\).
Result: 6/8 admission, zero type violations, versus 4/8 with the stage missing; only 5/8 versus a reordered-program control.
Epistemic status: abstained—useful typed stage construction, but insufficient identification of its internal composition.

The relative nature of the creativity claim matters. The B5 proposal is transformational relative to the task-stage vocabulary because it adds a new typed operator. It is combinational relative to the registered fragment grammar because it does not invent a new primitive or its contract. Because admission is withheld, however, the study does not establish a transformational result. This nested reference frame is more informative than labeling the episode simply creative or noncreative.

B5.1 therefore changes the experimental observer rather than enlarging the search. It actively chooses games or interventions that elicit edge-level counter-witnesses between candidate programs, then freeze and retest the surviving extension. Accommodation proceeds only after the system can distinguish programs that happen to agree on coarse task success.

B5.1 implements this observer. A deterministic scout visits a pool of 24 training worlds and, at each state, asks what action each of the 24 candidate programs would recommend. It never observes reward, terminal success, an expert plan, or the evaluator-only canonical program. The acquisition score counts pairwise action disagreement, states with disagreement, and the number of distinct recommendations. Eight maximally discriminating worlds are then selected under a frozen budget; a fixed eight-world sample supplies the passive control.

Active and passive acquisition produce different candidate programs inside the registered fragment grammar. The active sample selects

\[ \mathsf{Priority}(\mathsf{openUnvisited}, \mathsf{goNonDestination},\mathsf{goUnvisitedAny}), \]

which happens to equal the sequestered canonical program. The passive sample instead places \(\mathsf{goNonDestination}\) before \(\mathsf{openUnvisited}\). On twelve disjoint admission games, the active program solves 12/12, satisfies its registered contract on 12/12, incurs zero type violations, and requires 4.17 mean search steps. The passive program solves 11/12, satisfies its registered contract on 11/12, incurs 14 precondition violations, and requires 6.25 mean search steps. The active program wins their registered paired utility comparison by two games.

B5.1 nevertheless abstains for two distinct reasons. First, the missing-stage control stumbles to task completion on 11/12 games, despite satisfying no search contract and accumulating 128 precondition violations. Its scalar success is therefore only \(1/12\) below the selected program, short of the registered \(3/12\) margin. Second, the syntactic runner-up differs only in a third fallback that is never behaviorally decisive on the admission worlds. It is observationally equivalent to the selected program under all registered utilities, so the paired separation gate fails. Confirmation again remains sealed.

Experiment: RELIC–ALFWorld-B5.1. Acquisition: reward-blind disagreement scout over 24 worlds; eight actively selected versus eight passive worlds.
Active proposal: recovered the evaluator-only canonical fragment program without access to it.
Admission: 12/12 success, 12/12 contracts, zero violations, and 4.17 mean search steps; passive proposal 11/12, 14 violations, and 6.25 steps.
Decision: abstain—the malformed missing-stage baseline retained high scalar success, while the closest syntactic rival was behaviorally equivalent.
Lesson: candidate theories must be quotiented by observable behavior, and typed coherence cannot be subordinated to terminal reward.

The two localized gate failures determine the B5.2 protocol. It ranks equivalence classes of programs rather than raw syntax trees and compares the selected class with the best observably distinct rival. Its primary utility is lexicographic: contract and precondition coherence first, task reward second, and efficiency third. This protocol is registered as a new declaration rather than used to reinterpret B5.1 after observing its outcome.

Fresh reward-blind acquisition over 24 further training worlds selects eight counter-witness worlds. The 24 candidate syntax trees induce only eight observable behavior classes when quotiented by their constructed-stage action traces. The selected class contains six programs and is represented by

\[ \mathsf{Priority}(\mathsf{goNonDestination}, \mathsf{openUnvisited},\mathsf{goUnvisitedAny}). \]

Its strongest observably distinct rival begins instead with unrestricted navigation. A second member of the selected class is frozen to test whether proposal equivalence transports rather than being assumed.

B5.2 passes every admission gate. The selected representative solves 10/12 games, satisfies its registered contract on 12/12, incurs no invalid action or precondition violation, and requires 7.92 mean search steps. It wins the lexicographic paired comparison with the strongest distinct class on seven games and with the missing-stage control on all twelve. The second member of the selected class has no utility disagreement with the representative. By contrast, the missing-stage and type-invalid controls solve 7/12 but satisfy no contract and each accumulate 204 precondition violations.

Admission opens the sealed unseen confirmation partition. Both representatives of the selected class solve 8/8 games, satisfy their registered contracts on 8/8 games, incur zero violations, and require 11.75 mean search steps. The distinct rival also solves 8/8 coherently but requires 15.0 steps; the evaluator-only oracle requires 13.88. The missing-stage and type-invalid controls fall to 1/8 while each accumulates 201 violations. Thus the quotient preserves coherent behavior and an efficiency advantage under distribution shift, without claiming a unique syntax for the constructed operator.

Experiment: RELIC–ALFWorld-B5.2. Quotient: 24 syntax trees collapse to eight observed behavior classes; selected class size six.
Admission: 10/12 success, 12/12 contracts, zero violations; seven paired wins over the strongest distinct class and twelve over the missing stage.
Confirmation: 8/8 unseen success and contracts, zero violations; 11.75 mean search steps versus 15.0 for the distinct rival.
Epistemic status: admitted construction of a new typed task-stage operator up to an empirical behavioral equivalence class.
Boundary: primitives, the stage contract, and the admissible-command channel remain registered scaffolding.

Stronger tests must remove three kinds of scaffolding: invent or learn a new primitive or stage contract rather than only composing registered fragments; extend the state representation when the fragment grammar is insufficient; and learn the admissible-command model from interaction while preserving generated-action and validity controls. The resulting frozen learner must then be compared with capacity- and interaction-matched pretrained, fine-tuned, planning, and schema-learning baselines at benchmark scale.

This progression gives an operational empirical meaning to “learning by double involution.” Policy learning supplies local improvement. Structural accommodation supplies new decision objects. Creativity becomes defensible only when the system can construct such an object relative to a frozen reference language and when independent evidence admits it.

The claim remains realization-specific. These experiments implement typed side operations, behavioral proxies for mixed witnesses, finite-rollout observers, and admission gates; they do not prove that the ALFWorld controller is an internal double involution algebroid in an arbitrary tangent category.

CLIC, OPTIC, and RELIC isolate three different creative objects: causal structure, executable skill, and persistent decision structure. The next chapter asks whether their certified outputs can be exchanged and composed without erasing the distinctions among their evidence, uncertainty, and admission conditions.