ifc-0104

6.15.1 Biological foundation models, world models, and co-scientists

6.15.1 Biological foundation models, world models, and co-scientists

Current biological systems illustrate three components that a synthetic discovery architecture may combine. Evo 2 is a genomic foundation model trained on trillions of DNA base pairs across the domains of life. It predicts effects of genetic variation, represents biological features learned from sequence alone, and generates genome-scale sequences [ Brixi et al. , 2025 ] . It therefore supplies a powerful proposal distribution over a biologically meaningful sequence space, but generation within that space is not by itself an explicit theory of the mechanisms that make a sequence viable.

A complementary program seeks multimodal foundation models of human cells, connecting DNA, RNA, proteins, structures, images, literature, cell states, and higher biological scales. Biomni addresses the accompanying workflow problem: it combines language-model reasoning, retrieval, code execution, and a large environment of biomedical tools, databases, and protocols to carry out heterogeneous research tasks without a fixed task template [ Huang et al. , 2025 ] . A Stanford overview places Evo 2, virtual-cell modeling, and Biomni side by side as, respectively, a biological generative model, a prospective simulation substrate, and an agentic co-scientist [ Itoi , 2026 ] .

For synthetic creativity these roles should remain distinct. A domain foundation model proposes structured candidates; a virtual cell can act as a world model for simulated interventions; and an agent such as Biomni can orchestrate literature, databases, tools, and protocols. The theory layer must still declare what the system takes the biological objects and mechanisms to be, record which compositional obligation failed, and decide whether evidence from the simulator or laboratory admits a proposed extension. Together these systems make the experimental infrastructure far more powerful; they do not remove the need for an explicit and persistent object of scientific change.

The distinction is echoed by contemporary evaluations of idea generation. The Stanford report describes language-model proposals judged highly novel but often infeasible, with humans regaining the advantage when execution was included in the assessment [ Itoi , 2026 ] . In the terminology of this book, proposal novelty is not admission. A creative scientific system must connect an imaginative variation to executable intervention, evidence, and a theory package that can survive subsequent use.

The long-term target for synthetic creativity combines strengths from across this chapter: Palm’s context-relative account of novelty; Boden and Wiggins’s explicit conceptual spaces; Lenat’s and BACON’s heuristic discovery; the computational-creativity community’s process-sensitive evaluation; imagination machines as proposal generators; AI Feynman’s structural decomposition; and automated science’s experimental loop. Its distinctive wager is that these components can be coordinated around a persistent algebraic theory whose changes are explicit and whose models can be tested.

That claim remains an open empirical hypothesis. Neither a high novelty score, an impressive generated artifact, a recovered equation, nor a successful simulator intervention alone demonstrates field-building creativity. The achievement sought by this book is cumulative: a new presented language that organizes several results, survives independent criticism, and enables other agents to pose questions that were previously unavailable.

. Each experiment must state whether it evaluates artifact novelty, search, equation recovery, experimental discovery, or theory extension. Scores may be combined only after their semantics are recorded. In particular, predictive accuracy does not certify ontology change, and novelty does not certify scientific validity.

The landscape therefore supplies comparisons and evaluation instruments, not a single definition to be adopted wholesale. Its output is a set of questions that every experiment should answer: what changed, relative to which representation, by what mechanism, and under which certificate? The following table makes the handoff to the next three chapters precise. Their difference is not how impressive an output looks, but what remains fixed, what persistent artifact is produced, and what additional evidence that artifact owes.

Mode

What remains fixed

Persistent artifact

Admission burden

Combinational

source objects and their declared theories; an invented interface is recorded as an extension if it adds a type or law

a typed interface and reusable composite

legality, preserved obligations, nontrivial interaction, and transfer

Exploratory

the entire representational package and its legal moves

a path, construction, proof, policy, or consequence inside that package

validity, novelty relative to the registered space, independent value, and held-out consequence; search efficiency is reported separately

Transformational

no component is exempt, but every edit must be named and versioned

a versioned package comparison with transport, loss, and admission records; a theory extension when the presentation changes

failure of adequate in-language controls, independent discrimination, conservativity, and calibrated abstention

Table 6.3 The handoff from the literature landscape to Boden’s three modes. The modes classify changes relative to a registered representation; they are not rankings of historical importance.

The modes label segments of a versioned trace; they are not mutually exclusive classifications of whole systems. One episode may discover a typed interface, explore consequences made reachable by that interface, and later encounter evidence that forces a change of language. Boden’s mode and the episode’s Piagetian status are separate annotations: a combination may be assimilatory when its interface is already expressible, or proto-accommodative when it requires a proposed extension. The audit record should therefore mark where the maintained representation changes rather than assigning one creativity label retrospectively to the entire episode.

Chapters 79 now restate these modes in the common language of typed theories, infinitesimal probes, and admission. Part II then asks how those distinctions survive contact with executable DIAL algorithms.