ifc-0178

13.6 Competence of persistent structural artifacts

The A-series experiments evaluate frozen competence before adding online repair. A0 ran proposal-only, frozen-RELIC, and oracle-admissible policies on all 140 seen and 134 unseen TextWorld validation games. On unseen games, the proposal policy solved 0.201, frozen RELIC solved 0.910, and the oracle control solved 0.933. Relative to the matched proposal policy, frozen RELIC gained 0.709 success, reduced mean invalid-action rate from 0.708 to 0.083, and reduced median episode length from the 50-step ceiling to 21.

A1 tested simpler explanations. Uniform random generation and uniform random admissible choice both solved zero of 274 games. Across five unseen-house tie seeds, proposal-only success averaged 0.212 (standard deviation 0.034), while frozen RELIC averaged 0.918 (standard deviation 0.012). The structural gain was therefore not a favorable tie-breaking accident.

Stage

Question

Evidence

Base result

RELIC result

Inference

A0

Does a frozen artifact help?

140 seen, 134 unseen

0.201 unseen

0.910 unseen

Strong competence and transport.

A1

Is A0 a random or tie-seed effect?

All games; five OOD seeds

0 random; 0.212 proposal

0.918 mean OOD

Validity alone is insufficient; the gain is stable.

B1

Does generic novelty create structure?

24 training episodes; paired evaluation

0 success

0 success, positive shaped reward

Learning changed, but no task theory emerged.

B2

Do isolated structural proposals suffice?

Checkpoints at 12, 24, 36 episodes

0 success

0 success

One-step guidance did not compose.

B3

Can a macro-schema be admitted?

8 admission; 16 confirmation games

0 throughout

0.875 admission; 0.625 seen and unseen

A temporal composition can be admitted and transported.

B4

Can the composition be constructed?

24 orders on 6 proposal; 8 admission; 16 confirmation

0 DQN and shuffled

0.750 admission; 0.625 seen and unseen

The useful order is recovered, frozen, admitted, and transported.

Table 13.1 The RELIC–ALFWorld evidence sequence. A-series results test the value of persistent artifacts; B-series results distinguish shaped learning, isolated exploration, supplied structural accommodation, and combinational construction.

These are strong competence results, but competence is not the same as creativity. The schemas were already frozen when A0 and A1 began. They establish that structural artifacts can be useful, stable, and transportable; they do not establish how those artifacts were constructed.