ifc-0319

20.7 Evaluation beyond novelty

Novelty is cheap when the proposal space is large. Random corruption, adversarial optimization, and unconstrained sampling can all produce artifacts that are unlike a reference set. The more demanding question is whether a candidate extension earns a stable place in a theory. Across domains, at least eight tests recur:

  1. Structural necessity: does an independently measured obstruction motivate the extension?

  2. Irreducibility: is the proposal more than an obscure composition or renaming of admitted generators?

  3. Conservativity: does valid prior knowledge survive?

  4. Closure: does the proposal remove the targeted obstruction without creating larger unexplained failures?

  5. Productivity: does it generate a family of useful artifacts, predictions, or actions?

  6. Transport: does it work across tasks, contexts, models, or media not used to propose it?

  7. Independent value: do external observers find the extension scientifically, practically, or artistically worthwhile?

  8. Economy and safety: are the gains worth the acquisition, verification, and error costs?

These criteria turn evaluation from a contest of unusual outputs into an audit of theory change. They also make negative results informative. A system that abstains because the obstruction is transient, the extension is reducible, or the evidence does not transport has learned something about the boundary between novelty and warranted invention.