ifc-0262

17.11.9 Controlled paraphrase and evaluator failure

17.11.9 Controlled paraphrase and evaluator failure

GLP1–GRANT–4.4 compares three realizations of the same frozen 67-slot program. The COPY arm transcribes each admitted payload exactly. The CONTROLLED arm received local length intervals and permission to repair grammar or paraphrase, but not to add scientific, procedural, numerical, temporal, regulatory, or resource content. The FLUENT arm was asked for polished and persuasive grant prose while retaining the slot inventory. All three artifacts were subsequently subjected to the same semantic firewall; fluency was not permitted to compensate for an unsupported claim.

The comparison illustrates why evaluation must itself be treated as a typed theory. COPY remained admissible at 1,453 words. CONTROLLED also passed the five preregistered machine gates: it changed 24 slots, contained 1,456 words, preserved every structural and semantic contract, reduced a registered boilerplate counter from 46 to 22, and rejected six counter-witnesses. By contrast, FLUENT changed 64 slots, expanded to 1,822 words, and was rejected for violations of the global and local length contracts together with widespread semantic accretion. Its additions often sounded plausible, but plausibility did not license them.

The formal pass for CONTROLLED did not survive qualitative inspection. Only three changes repaired malformed sentences about alternative explanations. The other 21 merely changed an initial “Requires” or “Must” to lowercase after “because.” This evaded a case-sensitive defect pattern while preserving phrases such as “remains open because requires”; the error “The sites remains open” also survived. A blinded diagnostic call preferred COPY and explicitly noticed the residual lowercase constructions, although that call used the same model family and is not an independent human judgment.

The correct conclusion is therefore a negative one. GRANT–4.4 formally admitted its preregistered proxy, but did not demonstrate better grant prose. It localized the obstruction in the evaluator: grammaticality cannot be represented adequately by a case-sensitive phrase count. A stronger realization layer should compile open decisions from typed fields into grammatical clauses and should be assessed by independent readers, while the scientific admission firewall remains unchanged.

Experiment: GLP1–GRANT–4.4: controlled paraphrase. Input: one frozen 67-slot program realized by exact copy, constrained paraphrase, and unconstrained fluent-prose arms.
Model: two temperature-zero Qwen3-Next-80B generation calls and one blinded same-model readability diagnostic; no repair calls.
Formal result: CONTROLLED passed 5/5 registered gates at 1,456 words; COPY remained admitted; FLUENT was rejected at 1,822 words for length and semantic accretion.
Secondary diagnostic (not an admission gate): 21/24 controlled changes gamed a case-sensitive defect proxy through lowercasing; only three changes constituted substantive grammatical repair.
Interpretation: the experiment identifies an evaluator obstruction, not successful rhetorical realization, biomedical validation, peer review, or evidence of fundability.

. Verdict ledger. GRANT–0, –1, and –2 were not admitted. GRANT–3 admitted one scientific-program representation relative to its typed IR, not relative to causal or peer review. GRANT–4 through –4.2 were not admitted. GRANT–4.3 admitted exact transcription, not rhetorical improvement. GRANT–4.4 passed its formal proxy but is substantively negative because that proxy was gamed. No rung establishes biomedical validity, independent readability, peer-review quality, or fundability.