ifc-0206

15.12 An integrated benchmark

The preceding ladders isolate declaration, diagnosis, compilation, and active query selection. A full AGENTIC benchmark must test whether these capabilities compose rather than merely coexist. AGENTIC-MATH–0 begins with finite symmetry worlds. The evaluator withholds the number or arrangement of latent factors, a canonicalization operation, and the decision rule that would exploit the resulting orbit structure. The system receives exact transformations and bounded active queries. It must diagnose the missing invariant, construct an executable normal-form or quotient skill, and use that skill to select shorter discriminating query sequences in held-out worlds.

The full factorial comparison from Chapter 14 separates the structural component, OPTIC, RELIC, their pairs, and AGENTIC. Additional controls receive the hidden invariant, canonicalizer, or optimal query policy separately. Primary measures are exact structural recovery, constructor validity, query cost, false extension, and transport under renamed generators and changed factor counts. The central interaction question is whether the full system constructs a reusable primitive that no component or pair can both discover and exploit.

Experiment: AGENTIC-MATH–0. Epistemic status: proposed integrated benchmark, not a completed AGENTIC result.
World: generated finite symmetry systems with blinded names and withheld latent-factor organization.
Joint construction: invariant or quotient, executable canonicalization skill, and an adaptive query schema using it.
Admission: exact countermodels and held-out transformations, typed execution, improved query efficiency, and transport under renaming and changed factor count.
Controls: full component factorial, flat search, and separate structure, skill, and policy oracles.
Boundary: formal success establishes a result in the generated world, not historical-scale mathematical creativity.