ifc-0284

18.14.4 Strong generators change the opportunity

18.14.4 Strong generators change the opportunity

Outputs obtained through the OpenAI Playground show why the ARTISTIC claim must remain provider-conditioned. A terse request for a robot playing tennis produced a coherent scene with two opponents separated by the net, one racket each, and one ball. A similarly terse pool request produced a plausible mid-game shot. Under their generic visible-scene contracts, the correct action is admission without editing.

The same interface also produced more subtle failures under stronger contracts. A wide cricket request asking that every player and both umpires be visible generated fewer than the required eleven fielders, placed both umpires side by side at the bowler’s end rather than assigning one to square leg, and appeared to attach the ball to a rod or malformed object associated with an umpire. The nouns are present, but cardinality, spatial role, topology, and event provenance do not compose. A bases-loaded baseball image successfully placed three runners, yet assigned the catcher the blue uniform of the batting team while the pitcher wore the opposing red-and-white uniform. That failure is captured by the typed compatibility law

\[ \operatorname {team}(\mathrm{catcher}) =\operatorname {team}(\mathrm{pitcher}) \ne \operatorname {team}(\mathrm{batter\ and\ runners}). \]
Admit: coherent tennis scene Selective assurance on two OpenAI Playground outputs. The tennis image satisfies its visible-scene contract and should be left unchanged. The wide cricket image responds impressively to a demanding prompt but fails global player cardinality, umpire placement, and ball-source…
Repair required: global cricket contract Selective assurance on two OpenAI Playground outputs. The tennis image satisfies its visible-scene contract and should be left unchanged. The wide cricket image responds impressively to a demanding prompt but fails global player cardinality, umpire placement, and ball-source…
Figure 18.8 Selective assurance on two OpenAI Playground outputs. The tennis image satisfies its visible-scene contract and should be left unchanged. The wide cricket image responds impressively to a demanding prompt but fails global player cardinality, umpire placement, and ball-source constraints. The comparison illustrates two decisions by ARTISTIC, not a ranking of image providers.

Challenge case

Structural reading

ARTISTIC decision

Firefly cricket

Missing roles, incomplete pitch, foreign object, and ball scale defect

Global repair; retain after manual audit

Nano Banana cricket

Duplicate role and missing bowler

Rebind roles and reconstruct delivery

Nano Banana baseball

Incompatible action phases and excess agents

Select pre-pitch state; provisional admit

Nano Banana tennis

Opponents on same side; duplicate and embedded ball

Repair topology and contact geometry

OpenAI tennis and pool

Generic visible-scene obligations satisfied

Admit unchanged

OpenAI close cricket

Locally coherent; global roles outside the frame

Admit with global uncertainty

OpenAI wide cricket

Global count, umpire-position, and ball-source violations

Repair or regenerate

OpenAI baseball

Bases loaded, but catcher bound to the offensive team

Localized attribute repair

These observations are case studies rather than matched estimates of provider quality. Prompts, seeds, aspect ratios, sampling policies, model revisions, and editing interfaces were not controlled across systems. A proper benchmark must freeze those variables where interfaces permit and report both raw competence and marginal ARTISTIC gain:

\[ B_p=\operatorname {score}(I_p,S),\qquad \Delta _p=\operatorname {score}(I_p^+,S)-B_p, \]

where \(p\) indexes the provider and \(I_p^+\) is the retained output after any selective intervention. A strong base model may have large \(B_p\) and small \(\Delta _p\); ARTISTIC may still be valuable as a certificate for a private, high-cost, or exact domain contract. In casual generation, the same additional verification cost may not be justified.