sec-agentic-controlled-demonstration

14.10 A controlled demonstration: active evidence across the workflow

The first computational realization of the integrated workflow uses a small family of finite, partially observed worlds. A latent regime changes which of two actions is useful. The regime becomes observable only after a compound probe is executed, and the agent’s original decision language does not contain the response-conditioned policy needed to exploit the result. CLIC must support the latent distinction, OPTIC must establish the compound probe’s registered execution contract, and RELIC must construct and admit a policy that maps the probe’s two responses to different actions. The exact AGENTIC–0A ablation supplies the component-necessity result described above.

The evidence sequence replaces exact packets one at a time with finite-sample estimators; its separately registered rungs are not pooled. AGENTIC–0A verifies exact composition. AGENTIC–0B estimates CLIC and falls short; AGENTIC–0B.1 repairs the boundary cases through active acquisition. AGENTIC–0C localizes three OPTIC failures, and AGENTIC–0C.1 repairs them without relaxing the gate. AGENTIC–0D estimates RELIC. Because AGENTIC–0D.1 contains no eligible downstream boundary case, it does not test active acquisition; the fresh hierarchical AGENTIC–0D.2 split supplies that missing test.

In the final controlled-demonstration stage, all three component packets depend on empirical trajectories. CLIC actively gathers counter-witness evidence; an independent OPTIC check certifies the proposed probe’s registered execution contract; and RELIC estimates four reward cells—two probe responses crossed with two candidate actions. RELIC admits the new policy only when the preferred action in each response context has a Wilson interval separated from its alternative and the two preferences form a genuinely response-conditioned schema. Packet parsing and final execution remain exact. The workflow graph, packet schema, candidate probe, action vocabulary, and response-conditioned-policy form also remain registered. Thus the constructed object is a guarded decision extension inside a supplied meta-workflow; the experiment does not construct or revise the meta-workflow language itself.

For this experiment, overall workflow admission means that the CLIC packet passes its registered causal-evidence gate, the OPTIC packet passes its execution-contract gate, RELIC admits a differentiated response-conditioned policy, and both interfaces pass version, type, uncertainty, and provenance checks. This is a deliberately conjunctive finite certificate for the registered testbed; it is not a general theorem that arbitrary component certificates compose.

The RELIC test uses a preregistered mixture of standard worlds and boundary worlds. In the latter, the correct and incorrect actions have reward probabilities 0.64 and 0.36, rather than 0.70 and 0.30. Every response–action cell initially receives eighty trials. Active RELIC allocates batches of twenty additional trials to both actions of an unresolved response context, but only when the current point estimates already form a differentiated candidate policy. It is compared with a fixed budget, a uniform high budget of 120 trials per cell, and a matched-random control that permutes the active method’s realized extra budgets across worlds and response contexts.

Because an upstream abstention prevents RELIC from being evaluated, the primary endpoint is RELIC admission conditional on valid CLIC and OPTIC packets. Overall workflow admission is retained as a secondary endpoint. This hierarchical evaluation prevents a CLIC or OPTIC miss from being misreported as evidence against downstream acquisition.

Split

Acquisition

RELIC given upstream

Boundary RELIC

Trials per challenge

Admission

Fixed (80)

0.812

0.625

320.0

 

Matched random

0.812

0.625

329.2

 

Uniform (120)

0.938

0.875

480.0

 

Active RELIC

0.979

0.958

329.2

Confirmation

Fixed (80)

0.771

0.542

320.0

 

Matched random

0.792

0.583

334.2

 

Uniform (120)

0.958

0.917

480.0

 

Active RELIC

0.979

0.958

334.2

Table 14.2 AGENTIC controlled-demonstration results. Active RELIC targets unresolved response–action comparisons and is evaluated conditionally on valid upstream packets. The matched-random control receives the same realized additional sample budget. Admission and confirmation use disjoint registered surface families and random seeds.

Active acquisition repaired eight of nine eligible fixed-budget failures in the admission split and ten of eleven in confirmation. Matched-random allocation repaired none and one, respectively. The active method also slightly exceeded uniform acquisition while using approximately 30% fewer decision trajectories. Uncertainty-gated RELIC admitted no policy under shuffled rewards and no differentiated policy in forced-null worlds; the corresponding point-estimate control admitted 25.0% and 18.8% of shuffled cases. Conditional interval coverage was 98.9% in both prospective splits.

The two remaining abstentions illustrate the guard rather than an exhausted budget. In admission, one unresolved response had empirical margin 0.075, below the registered 0.08 boundary required to justify further sampling. In confirmation, the provisional policy assigned the same action to both responses and therefore did not constitute a differentiated extension. Active RELIC declined to spend additional evidence in both cases.

. This experiment is a controlled mechanism demonstration, not evidence of general scientific creativity. The worlds are synthetic; the probe and action vocabularies are registered; the proposed extension has a known two-response, two-action form; reward noise is Bernoulli; and packet parsing and deployment are exact. The bounded construction claim tested here is the guarded admission of a response-conditioned policy absent from the baseline declaration. It does not include invention of new actions, sensors, objects, or an unrestricted theory language.

Evidence beyond this mechanism demonstration would have to relax these conveniences separately: hide or generate the candidate vocabulary, enlarge response and action spaces, introduce nonstationary and partial observation, learn packet interfaces, or move to language-mediated and physical environments. Within the registered synthetic workflow, the present result establishes only that the composed architecture can localize an evidential obstruction, acquire data at that obstruction, and retain conservative abstention when its declared preconditions are absent.