ifc-0254

17.11.1 Structural compilation without semantic admission

17.11.1 Structural compilation without semantic admission

GLP1–GRANT–0 asks a deliberately narrower question than whether an AI can write a fundable proposal: does the maintained structure accumulated by Democritus and GLP1–PILOT–1 improve the construction of an auditable research program? The controlled comparison used three temperature-zero generations from the same GPT-OSS 20B model and the same frozen five-study corpus. The flat condition received ordinary article summaries; ledger received provenance-bearing causal claims and restriction obligations; and compiled additionally received the exposure-history obstruction, a three-aim program grammar, dependency rules, a mock funding call, and an explicit team-and-resource profile.

All conditions received the same output schema. The registered score measured ten aspects of contract satisfaction, including program- and aim-field coverage, citation validity, per-aim provenance, controls, alternative outcomes, discriminating observations, dependency validity, context alignment, resources, and epistemic boundaries. It did not score biological importance, novelty, ethics, feasibility, or fundability.

Condition

Score

Program

Aims

Valid sources

Provenance

Flat

0.892

1.000

0.917

1.000

1.000

Ledger

0.892

1.000

0.917

1.000

1.000

Compiled

1.000

1.000

1.000

1.000

1.000

The result was not admitted: eleven of twelve preregistered gates passed. The compiled artifact improved on the flat condition by 0.108, short of the frozen 0.15 margin. This ceiling effect is itself diagnostic. Once every condition is given a detailed JSON schema, a capable language model can produce a proposal-shaped document with valid citations, controls, alternatives, and a dependency graph. The compiled representation supplied the missing exposure-history and resource alignment, but structural completeness alone did not sharply separate it from ordinary generation.

A retrospective semantic audit, conducted outside the registered gates, revealed the more important limitation. The compiled program invented participant counts, trial durations, doses, blinding choices, endpoints, and claims of resource sufficiency that were not licensed by the ledger. It treated access to five source summaries and human-curated arm events as if participant-level trial records were available, and it proposed transport to a cardiovascular population without a sufficiently grounded bridge. Its experiments largely compared efficacy rather than discriminating among mechanistic explanations of regain and durability. The artifact therefore passed the structural contract while remaining scientifically premature.

This negative result identifies a missing layer. A grant constructor needs a semantic admission firewall in addition to a document schema. Every protocol commitment must be classified as source-supported, externally supplied, or hypothetical; unsupported numerical and operational choices must remain open variables; resource feasibility must be checked against actual capacity; and experiments must be scored by whether they distinguish declared causal alternatives. GLP1–GRANT–0 thus supplies a useful baseline and exposes why grant construction cannot be reduced to well-formed scientific prose.

Experiment: GLP1–GRANT–0: controlled program compilation. Input: five frozen GLP-1 study summaries, with flat, claim-ledger, and fully compiled representations.
Model: GPT-OSS 20B, one temperature-zero generation per condition.
Result: compiled score 1.000 versus 0.892 for flat and ledger; 11/12 registered gates; not admitted because the 0.15 improvement margin was not met.
Semantic audit: unsupported protocol and resource commitments remain despite perfect structural contract satisfaction.
Boundary: one small retrospective compilation test, not expert peer review, scientific validation, clinical guidance, or evidence of fundability.