ifc-0319
20.7 Evaluation beyond novelty
Novelty is cheap when the proposal space is large. Random corruption, adversarial optimization, and unconstrained sampling can all produce artifacts that are unlike a reference set. The more demanding question is whether a candidate extension earns a stable place in a theory. Across domains, at least eight tests recur:
Structural necessity: does an independently measured obstruction motivate the extension?
Irreducibility: is the proposal more than an obscure composition or renaming of admitted generators?
Conservativity: does valid prior knowledge survive?
Closure: does the proposal remove the targeted obstruction without creating larger unexplained failures?
Productivity: does it generate a family of useful artifacts, predictions, or actions?
Transport: does it work across tasks, contexts, models, or media not used to propose it?
Independent value: do external observers find the extension scientifically, practically, or artistically worthwhile?
Economy and safety: are the gains worth the acquisition, verification, and error costs?
These criteria turn evaluation from a contest of unusual outputs into an audit of theory change. They also make negative results informative. A system that abstains because the obstruction is transient, the extension is reducible, or the evidence does not transport has learned something about the boundary between novelty and warranted invention.