ifc-0133
9.8 Evaluating transformational creativity
Predictive accuracy on the motivating data is necessary but weak: a more expressive theory can often fit what exposed the failure. Evaluation should instead form a vector of criteria:
- Resolution
Does the extension remove the localized obstruction?
- Minimality
Could a less disruptive change do the same work?
- Conservativity
Which earlier results and skills survive?
- Novel consequence
Does the extension entail a test not used to propose it?
- Discrimination
Can an experiment separate it from rival extensions?
- Unification
Does it organize several failures through one construction?
- Persistence
Is the new object or operation reused in later tasks?
- Calibration
Does the system defer admission when evidence is inadequate?
No scalar combination of these criteria is canonical. The vector should remain visible so that a compact but destructive theory is not silently ranked above a conservative one, or a novel vocabulary above an empirically grounded extension. Unification may reduce description length, but compression alone does not discharge the other admission obligations.
. Admit a transformation only with a versioned package comparison—including a theory map when the presentation changes—a localized obstruction, comparison against in-language repairs, transport account, independent consequence, and explicit epistemic status. A fluent proposal without these components remains a hypothesis, not an admitted creative jump.