sec-relic-b3

13.8 Admitting a supplied temporal composition (B3)

B3 freezes the episode-36 structural TextDQN from B2 and registers the candidate

\[ \mathsf{LocateTarget} \; \to \; \mathsf{AcquireTarget} \; \to \; \mathsf{LocateDestination} \; \to \; \mathsf{PlaceTarget}. \]

The schema consumes only the natural-language task, observations, action history, and the same admissible-command channel available to the matched DQN. It activates only after the base policy repeats the same action three times. This repetition is the registered obstruction: a local policy loop that makes no progress toward the declared task.

The evidence is divided before execution. Four discovery games document that the frozen DQN repeatedly exhibits the obstruction. Eight disjoint seen games compare the unchanged DQN and candidate schema for admission. Admission requires at least 0.75 candidate success, at least 0.50 absolute gain over the DQN, at least 0.75 trigger coverage, and zero invalid actions. Only an admitted schema may proceed to eight untouched seen and eight untouched unseen games.

The frozen DQN solves none of the discovery, admission, or confirmation games. The candidate activates on all eight admission games, generates no invalid action, and solves seven. It therefore passes every admission gate. Once frozen, it solves five of eight untouched seen games and five of eight unseen games, while the DQN remains at zero in both splits. The Wilson 95-percent interval is \([0.529,0.978]\) for admission and \([0.306,0.863]\) for each confirmation split.

This is the separately registered 30-step horizon-recovery record. An earlier 15-step execution was invalidated before confirmation and is not counted as an additional trial.

Experiment: RELIC–ALFWorld-B3. Target and obstruction: complete simple object placement; three consecutive identical actions without completion.
Candidate: a four-stage typed pickup-and-place composition.
Evidence: four discovery games; eight admission games; eight seen and eight unseen confirmation games.
Gates: locked success, gain, trigger, and validity thresholds.
Result: 7/8 admission; 5/8 seen and 5/8 unseen confirmation; zero invalid candidate actions.
Epistemic status: admitted supplied policy mechanism; experimenter-proposed schema.

The negative and positive rungs belong together. Had B3 followed only A0, its gain could be dismissed as a hand-coded planner. B1 and B2 show why the macro is structurally meaningful: dense novelty and isolated task-relevant actions both change learning, but neither supplies the missing composition. B3 adds exactly that object and subjects it to independent admission.