ifc-0263

17.11.10 From an AI-authored study to an AI-constructed program

17.11.10 From an AI-authored study to an AI-constructed program

The Open Conference of AI Agents for Science provides a useful contemporary contrast. Its inaugural call asked AI systems to lead hypothesis generation, experimentation, and manuscript writing, and required an AI system to be the sole first author. Submitted papers were evaluated for familiar research-paper qualities including quality, clarity, significance, and originality, with explicit disclosure of the relative AI and human contributions [ Open Conference of AI Agents for Science , 2025 ] . This is an important test of whether an agent can produce a credible completed unit of research and report it in the conventional scientific form.

A complementary benchmark from Inherent Laboratories makes replication itself the training target. Its Replica task space contains 310 tasks drawn from 100 machine-learning and AI-for-science papers. An agent must reconstruct a reported figure under bounded time and compute, without access to the original plot. The authors train a 27-billion-parameter agent, Faraday, by long-horizon reinforcement learning with coding agents as tools, and report stronger held-out replication performance than the frontier-agent baselines in their study [ Falck et al. , 2026 ] . This is more than document imitation: underspecified implementation choices must be recovered through hypothesis-driven experimentation.

Faraday nevertheless occupies a different point in the scientific workflow from the capstone developed here. Its registered artifact is a faithful reconstruction of an existing empirical result. GLP1–GRANT instead asks for a prospective theory of which unresolved claims deserve new evidence, how several aims should depend on one another, and why the proposed program merits scarce resources. A mature system could compose the two capabilities: theory construction proposes and justifies an experiment; a replication-trained agent executes or audits the computational protocol; the returned evidence then drives another round of theory repair.

The GLP1–GRANT capstone shifts the target from a completed study to a prospective research program. The central question is not only whether an agent can select a tractable hypothesis, execute experiments, and narrate the results. It is whether the agent can identify a consequential unresolved question, organize multiple dependent aims around it, anticipate results that do not yet exist, and justify committing scarce human and material resources to the program.

AI-authored research paper

AI-constructed research program

Reports completed experiments

Justifies experiments not yet performed

Demonstrates a result

Identifies a consequential unresolved question

Optimizes one investigation

Coordinates several dependent aims

Evaluates an artifact retrospectively

Evaluates importance, feasibility, and risk prospectively

Provides evidence for a claim

Argues why acquiring particular evidence merits investment

These are complementary tests of AI for science. The distinction concerns the target artifact, not a hierarchy of scientific value.

The two capabilities are complementary rather than competing. An admitted research program could hand its experimental aims to a Robot Scientist or an Agents4Science-style execution system. The resulting measurements and papers would then return as evidence for the next round of theory repair. This closes a larger loop from literature, through prospective program construction, to experiment, publication, and revised theory. In compact form, Agents4Science asks whether an AI agent can conduct and report a study; the GLP1–GRANT capstone asks whether it can construct a credible future course for a field.