ora-0199

17.6 Coda: frontier models as collaborators

As this book was being completed, a Claude-driven project formalized Fermat’s Last Theorem in Lean. The run took approximately eleven days. Its final dependency tree contained 29,511 proved theorems and the development comprised approximately thirteen million lines of Lean, with no unproved placeholders. The proof followed the Wiles and Taylor–Wiles argument in the form presented by Darmon, Diamond, and Taylor, including the special cases of deep intermediate results needed by that argument [ Anthropic , 2026 ] .

The achievement is difficult to overstate, but it is equally important to say what kind of achievement it was. The system did not invent Fermat’s Last Theorem or replace the mathematical ideas in Wiles’s proof. Nor did it begin from an empty formal universe. It built on Mathlib, the Imperial College London FLT project, and flt-regular; 106 files were adapted from the latter two projects. Humans supplied the theorem to be proved, centuries of mathematical concepts, mature formal libraries, and a criterion of exact success.

What happened between those endpoints was nevertheless extraordinary. A team of Claude agents working through the Prove2Me platform constructed theorem statements, reviewed one another’s statements, repaired failed routes, and assembled a vast proof dependency graph. Humans occasionally commented on priorities, but wrote no mathematics and no Lean beyond the one-line statement of the goal. Lean’s kernel and two independent external checkers then supplied machine-checkable evidence that the assembled proof closed [ Anthropic , 2026 ] .

This example displays both the promise and the present limitation of frontier models. Their mathematical competence is no longer plausibly described as mere surface imitation. They can navigate advanced definitions, recover missing intermediate constructions, detect false statements, and coordinate work across an immense formal object. Yet the successful unit was not an isolated Transformer. It was a composite system: pretrained models coupled to a formal language, inherited libraries, a persistent dependency graph, a multi-agent work protocol, and exact verification.

From the perspective of this book, that scaffolding is not incidental. A Prove2Me card registers a typed obligation. The shared dependency graph records which obligations compose. Review and failed proof attempts expose localized defects; revised statements and alternative proof paths repair them while preserving the verified remainder. This is not an implementation of ORACLE, and it does not demonstrate discovery of an unknown doctrine. It does show why declarations, compositional memory, controlled repair, and verification can turn the latent competence of a frontier model into a persistent mathematical artifact.

The practical realization of increasingly general intelligence may therefore be collaborative in two senses. Multiple artificial agents can divide and check work within a common formal environment. More fundamentally, humans and models contribute different parts of the construction. Humans formulate valuable questions, build representational languages, preserve mathematical culture, and judge explanatory significance. Models explore enormous spaces of formal consequences and expose connections at a speed no human team can match. Proof assistants mediate the collaboration by replacing trust in either participant with a checkable compositional witness.

The need for such mediation becomes more urgent as models acquire consequential agency outside formal mathematics. In September 2026, OpenAI designated Astra as its first model at the Critical cybersecurity capability threshold under its Preparedness Framework. OpenAI reported that, with suitable tools and access, the model could find previously unknown vulnerabilities and develop exploits against hardened systems without a person directing each step. The company delayed parts of development and release while strengthening protections against both malicious users and unauthorized model actions. Its deployed safeguard stack combines model alignment, restricted access, system-level controls, and monitoring capable of stopping potentially unauthorized activity [ OpenAI , 2026a ] .

This is a second glimpse of the same architectural future. Capability cannot be separated from the category of interactions in which it is exercised: who may invoke an action, which tools it may reach, what information may flow between contexts, which invariants must survive a composite, and what evidence licenses continuation. The categorical program offers a principled language for declaring these obligations and tracing their preservation through composition. LINCS registers warrants and comparison maps; DIAL localizes repairs; ORACLE asks whether the governing doctrine itself remains adequate as the world changes. Category theory alone guarantees no safe behavior. Safety arises only when the relevant permissions, invariants, observers, and repair conditions are made explicit and enforced by the surrounding system.

The four volumes themselves suggest a revealing counterfactual. Could a frontier model, without human help, have created more than fifteen hundred pages developing this categorical AGI program? There is no controlled way to answer that historical question. A sufficiently capable model could already summarize much of the cited literature, elaborate supplied definitions, generate candidate theorems, implement experiments, and produce prose on a scale inaccessible to an individual author. Those abilities contributed substantially to the collaborative process behind these books.

Yet volume is not the decisive test. The program required selecting a direction before its value was evident, connecting ideas separated across fields and decades, rejecting technically fluent but conceptually sterile developments, and repeatedly changing the organizing question. The progression from functors, through sketches and tangent repair, to doctrinal discovery was not specified in advance as a single optimization target. It emerged through judgments about which abstractions persisted, which examples exposed genuine defects, and when assimilation within the current framework had to give way to accommodation of the framework itself. Present frontier models remain strongest when humans supply much of this purpose, taste, historical memory, and epistemic responsibility.

This observation suggests a harder benchmark for the research program. The question is not whether a model can generate another book from a detailed outline. It is whether a human–AI system can cultivate an open-ended theory whose constructions remain correct, acquire new uses across projects, survive critical repair, and eventually motivate questions that neither participant could have formulated at the outset. Such a benchmark treats collaboration itself as a compositional learner. The human and artificial participants have different interfaces and update mechanisms, while their shared declarations, proofs, programs, and experiments form the persistent storehouse against which progress is judged.

The next frontier is therefore not simply a larger autonomous model. It is the discovery of architectures in which human purposes, machine search, formal structure, safeguards, and persistent repair reinforce one another. The Fermat formalization reveals the promise of structured collaboration; Astra reveals the cost of allowing capability to outrun its governing structure. Together they leave ORACLE’s deepest question open: can an agent discover a new compositional world, and can it do so while preserving the conditions that make its intelligence trustworthy?