ifc-0291
18.17 Measurements and failure modes
Pixel similarity is inadequate because the same semantic construction may have many visual realizations. The benchmark therefore uses typed scene-graph checks in the procedural phase, registered vision models and human audits in the natural-image phase, and controlled contrastive or perturbation-based tests where the protocol supports them. These tests diagnose structural dependence; they acquire a causal interpretation only when the manipulated interface has a registered causal semantics.
Primary measures include semantic accuracy, exact relation and binding success, preservation of irrelevant attributes, productivity on unseen combinations, sample and compute cost, false extension, abstention, and transport. The three-way AGENTIC interaction asks whether joint structural diagnosis, executable language construction, and active probe selection exceed all component and pairwise systems under matched budgets.
Characteristic failures include a token that memorizes its exemplars, a relation that works only for one object pair, an adapter that improves visual quality while changing the proposed meaning, a private language that no fresh agent can acquire, and a false accommodation caused by limitations of the visual evaluator. Each failure should remain in the theory store as a typed counterexample.