ifc-0103

6.15 Large language models in scientific discovery

Moving from law recovery to theory building requires evidence that can answer back. Automated laboratories provide one route because they select and execute experiments rather than merely fit a static dataset. The Robot Scientist “Adam,” for example, coupled hypothesis formation, experiment selection, laboratory automation, and interpretation in a closed functional-genomics setting [ King et al. , 2009 ] . Such systems make admission an active process constrained by instruments, time, cost, and biological variability.

The rapid growth of language-model systems for science supplies a second contemporary comparison. Zheng and colleagues organize this literature by an autonomy axis: an LLM may serve as a tool for one stage of inquiry, an analyst that integrates information and proposes conclusions, or a scientist that orchestrates substantial portions of the research cycle [ Zheng et al. , 2025 ] . Their accompanying, continuously updated bibliography spans literature synthesis, hypothesis generation, experimental planning, data analysis, validation, function discovery, autonomous research agents, and scientific foundation models [ HKUST KnowComp , 2026 ] . This is a valuable operational map of a fast-moving field.

Autonomy, however, is not the same axis as structural creativity. A system can autonomously traverse a hypothesis space, operate tools, and optimize an experimental objective while leaving the objects, admissible operations, and laws of its scientific language unchanged. Conversely, a human-guided system may propose a consequential change to that language. Synthetic creativity therefore adds a theory-change axis: whether a system searches within a fixed model class, recovers latent structure, or constructs and admits an extension of the theory itself.

In Piagetian terms, autonomy measures how much of the research loop the system executes without human direction; it does not measure whether the loop is assimilatory or accommodative. A highly autonomous agent may repeatedly assimilate evidence into a fixed ontology, while a human-guided collaboration may perform the decisive accommodation by introducing a new scientific object or mechanism.

The Scientific Discovery Evaluation (SDE) framework makes this distinction empirically concrete. It evaluates frontier models on 1,125 expert-vetted questions drawn from 43 research scenarios in biology, chemistry, materials, and physics, and separately on eight closed-loop projects in which models propose hypotheses, receive simulator or experimental feedback, and revise their proposals [ Song et al. , 2025 ] . The study finds a substantial gap between ordinary science-question performance and discovery-grounded performance, diminishing returns from model scaling and test-time reasoning, correlated failure modes across leading models, and no single model that dominates all projects. It also reports cases of successful optimization despite weak scenario-level knowledge, suggesting that guided or serendipitous exploration can sometimes compensate for an incomplete explicit account of the governing structure.

The authors also delimit the benchmark: its 43 scenarios and eight projects reflect a finite contributor cohort, omit several scientific and engineering fields, and use one evolutionary search strategy and prompting protocol for the more expensive project-level evaluations. Commercial API variation and the restriction of those evaluations to a subset of models further limit reproducibility and breadth. SDE is therefore a strong, explicitly bounded baseline rather than a universal assay of scientific discovery [ Song et al. , 2025 ] .

These results make SDE a strong baseline for this book, but its closed-loop projects generally begin with a supplied hypothesis space, computational oracle or simulator, and selection rule. The principal output is a successful candidate within that registered world. Our experiments ask a complementary question: can a system recover the world’s latent factors, generators, relations, invariants, and domain of validity, and can it enlarge that presentation when the registered language is inadequate? High objective value and accurate answers remain evidence, but neither alone certifies the construction of a reusable theory.