ifc-0243

17.8.1 Source discipline (CTTE–0)

17.8.1 Source discipline (CTTE–0)

Before confronting paper-length prose, entity aliases, and incomplete search, CTTE–0 isolates one question: once source claims have been extracted correctly, can a contextual theory constructor use them to recover a missing factorization and derive a relation that no source in its construction set states directly? It is a unit test of that operation, not a synthetic scientific domain or a benchmark of literature understanding.

The experimental units are 600 independently seeded claim ledgers—200 from each of three structural families. A ledger is a documentary projection of a hidden benchmark model. It contains 20 or 60 one-sentence records; each record carries a source identifier, corpus split, named context, evidence-type tag, and one signed perturbational claim, such as

Evidence type: controlled. In context A, perturbing stimulus_37 increased mediator_37.

The hidden model is the answer key used to generate and score the records; it is never given to the constructor. It specifies a registered intervention \(X\), outcome \(Y\), contexts \(A\) through \(D\), and either no mediator, one shared mediator, or incompatible context-specific mediators. The generator emits five separately identified reports for each relation and reverses each reported sign independently with probability \(0.05\). It then partitions source IDs into discovery, admission, and confirmation sets.

hidden benchmark model \(\longrightarrow \) noisy source reports \(\longrightarrow \) claim ledger \(\longrightarrow \) proposed sketch \(\longrightarrow \) sealed-report test

The ledger therefore defines what documentary evidence is available to the constructor, not what is true in a physical world. The hidden model defines truth only inside the benchmark. CTTE–0 supplies parsing, stable entity names, contexts, and the allowed one-mediator grammar so that failures can be attributed to contextual composition rather than natural-language processing.

The initial ontology contains \(X\) and \(Y\) but not a mediator. Three corpus families test the structural decision. In closed corpora, \(X\to Y\) is invariant and no extension is needed. In positive corpora, one undeclared object \(Z\) appears in every context: the sign of \(X\to Z\) changes with context, \(Z\to Y\) remains invariant, and composition explains the corresponding change in \(X\to Y\). In unsupported corpora, each context has a different mediator, so no one-object extension can unify the literature.

The target-context arrows \(X\to Y\) are absent from discovery and admission. They are opened only after the theory and two target predictions are frozen. The contextual constructor admits the factorization \(X\longrightarrow Z\longrightarrow Y\) only when one source-backed object explains both discovery contexts and its incident arrows survive admission in both target contexts. The proposal retains the local signs rather than averaging them.

We compare four representations. Relation extraction and local graphs cannot predict an unreported target edge. A global graph aggregates the two opposing discovery effects and abstains. The contextual sketch can compose through \(Z\). These zero-coverage baselines are intentional: CTTE–0 tests whether the locked relation is derivable from a shared contextual theory, not whether a larger predictor can guess its label.

This makes the scope of the result deliberately narrow. CTTE–0 asks whether the proposed representation supports source-preserving derivation, correct no-change decisions, and abstention when its extension language is inadequate. It does not test paper reading, entity resolution, open-ended mediator invention, experimental design, or causal truth. CTTE–1 through –5 isolate selected pieces of those harder capabilities under generated-world controls; paper understanding and causal truth remain outside the ladder and ultimately require the external admission loop described above.

Across 200 corpora from each family, structural action accuracy is \(99.7\% \). The system constructs and exactly recovers \(Z\) in \(99.5\% \) of positive worlds, retains the old theory in \(99.5\% \) of closed worlds, never falsely extends a closed world, and abstains in every unsupported world. It predicts both sealed target relations in \(99.5\% \) of positive worlds. Every admitted theory element retains a valid source identifier, all contextual contradictions are preserved, and no source identifier crosses a corpus split.

Representation

Target coverage

Accuracy

Exact-pair power

Relation extraction

\(0\% \)

\(0\% \)

\(0\% \)

One global graph

\(0\% \)

\(0\% \)

\(0\% \)

Context-local graphs

\(0\% \)

\(0\% \)

\(0\% \)

Contextual sketch extension

\(99.5\% \)

\(99.5\% \)

\(99.5\% \)

Experiment: CTTE–0: source-ledger theory construction. Input: 600 independently seeded synthetic claim ledgers: 200 per family, with 20 or 60 one-claim records and disjoint discovery, admission, and confirmation sources.
Withheld structure: one contextual mediator \(Z\) and two target relations.
Structural result: \(99.7\% \) action accuracy; \(99.5\% \) construction and mediator recovery.
Controls: \(99.5\% \) closed specificity, zero false extension, and \(100\% \) unsupported abstention.
Confirmation: \(99.5\% \) exact two-context prediction with complete provenance and contradiction preservation.
Verdict: admitted as a source-preserving derivability unit test with a supplied parser and one-mediator grammar, not as scientific discovery.