ifc-0236

17.3 Documents as implicit sketch presentations

The categorical form of this task is a document-to-sketch construction. A scientific document collection \(\mathsf D\) rarely declares a category, a path equation, or a limit cone. It nevertheless presents them implicitly. Noun phrases propose objects; perturbations and measurements propose arrows; experimental conditions type local contexts; statements of invariance, mediation, or equivalence propose equations and factorizations. Constructing a theory means making this latent presentation explicit.

The passage is staged rather than magical:

\(\mathsf D\xrightarrow {\; E_{\theta }\; }P_{\mathsf D} \xrightarrow {\; R\; }S_{\mathsf D} \xrightarrow {\; i\; }S_{\mathsf D}^{+} \xrightarrow {\; \operatorname {Adm}\; }\mathsf{Test}_{\mathsf D}\).

Here \(E_{\theta }\) is a versioned extractor, parameterized by an ontology and language model, that converts source spans into a provenance-bearing presentation \(P_{\mathsf D}\). The presentation contains candidate objects, arrows, contexts, equations, evidence types, and links back to the exact source spans. The reconciliation map \(R\) resolves aliases and incompatible contexts to form a source-bearing sketch \(S_{\mathsf D}\). A model of the extracted theory in a semantic category \(\mathcal{C}\) is then a sketch model \(M_{\mathsf D}:S_{\mathsf D}\longrightarrow \mathcal{C}\) that realizes the declared cones and equations. When the local presentations cannot be glued, the obstruction may justify an extension \(i:S_{\mathsf D}\to S_{\mathsf D}^{+}\): a new object, context split, factorization, equation, or experimental interface. Finally \(\operatorname {Adm}\) compiles distinguishing consequences into an admission interface \(\mathsf{Test}_{\mathsf D}\)—a laboratory protocol, simulator, observation campaign, physical model, or locked evidence query.

This notation should not be read as asserting a unique functor from prose to theory. Extraction and reconciliation can be ambiguous. More faithfully, a compiler returns a family \(\mathsf{Compile}_{\theta }(\mathsf D)=\{ (S_k,w_k,\pi _k)\} _{k\in K}\), where \(S_k\) is a candidate sketch, \(w_k\) records its support or uncertainty, and \(\pi _k\) maps every generator and relation to its provenance. A corpus inclusion \(\mathsf D\subseteq \mathsf D'\) should induce a conservative sketch map only when the new evidence preserves the old typing and identifications. If a new source forces two objects to split, two aliases to merge, or an equation to be withdrawn, theory revision is not a monotone database insertion; the old and new sketches require an explicit versioned transport.

This is also where double involution enters. Assimilation updates the evidence and models of a fixed \(S_{\mathsf D}\). Proto-accommodation proposes a changed presentation; only an independent admission process may admit \(S_{\mathsf D}^{+}\). Their interaction asks whether revising the interpretation of the documents and revising the theory commute. A persistent failure of the two orders to agree localizes the creative frontier: the documents cannot be coherently represented inside the current theory and extraction language.

CTTE–0 begins after the first arrow. Its input is already a parsed provenance-bearing presentation \(P_{\mathsf D}\), and it tests whether contextual reconciliation and finite extension derive locked consequences. CTTE–1 and CTTE–2 progressively move leftward by hiding entity registration and context boundaries. The full GLP-1 challenge must eventually evaluate the entire chain from paper-length documents through an externally executable admission interface.