ifc-0287
18.14.7 Three practical axes
18.14.7 Three practical axes
The cross-provider benchmark must measure more than final correctness. Generality asks whether the same engine and constraint vocabulary transfer across providers, domains, styles, action phases, and failure types. Ease of use measures human authoring time, domain expertise, annotations, correction turns, and whether a rulebook or manual can replace a hand-written sketch. Processing cost measures generation and editing calls, latency, compute, image tokens or pixels, and monetary cost. A useful summary quantity is the gain in verified structure per unit of additional burden,
with weights fixed by the intended application rather than selected after the result is known.
The matched evaluation should compare raw generation, sketch-compiled generation, one global repair, localized masked repair, and—where available— diffusion-native enforcement. Primary endpoints are constraint satisfaction, false repair of valid images, preservation of unaffected content, calibrated abstention, provider and domain transfer, latency, and cost. The strongest practical result may be selective rather than maximal: a terse request is generated once, most valid images are certified cheaply, and only supported violations trigger another expensive call.
. ARTISTIC scales by changing the theory supplied to a common assurance engine, not by hiding a separate hand-written repair inside every image. A better generator should reduce the intervention rate; it should not change what the registered visual contract means.