tab-dial-url2

15.13.9 Active coalgebraic discrimination

15.13.9 Active coalgebraic discrimination

Passive evidence may leave two behavior functors observationally equivalent even when a legal intervention would separate them. DIAL-URL–2 therefore turns model selection into sequential experimental design. Its finite probe menu contains neutral, observation-sensitive, duration-sensitive, and joint interventions. Each probe specifies an initial-condition family, an action schedule, and an observation protocol. The active learner evaluates the expected posterior entropy reduction under every probe and executes the most informative one. A matched passive learner instead cycles evenly through the same menu. Both learners receive the same prior, likelihood, observations, confidence threshold, and total budget.

The registered hypotheses are again MDP, POMDP, SMDP, and partially observable SMDP, represented by the constructor sets

\[ \varnothing ,\qquad \{ \mathsf{Obs}\} ,\qquad \{ \mathsf{Time}\} ,\qquad \{ \mathsf{Obs},\mathsf{Time}\} . \]

Every trajectory probe returns two observable counter-witness summaries: whether the registered history dependence appeared and whether a non-unit holding time appeared. A type is admitted only when its posterior probability reaches 0.95; otherwise the system abstains. Five hundred independently seeded replications were run for every true type and acquisition policy. Table 15.13.9 compares fixed-budget behavior.

Probes

Passive: exact / wrong / abstain

Active: exact / wrong / abstain

Exact gain

4

4.75% / 0.05% / 95.20%

20.30% / 0.70% / 79.00%

+15.55 pp

8

52.10% / 1.20% / 46.70%

75.45% / 1.00% / 23.55%

+23.35 pp

12

69.60% / 0.60% / 29.80%

92.85% / 0.40% / 6.75%

+23.25 pp

20

91.30% / 0.30% / 8.40%

99.65% / 0.00% / 0.35%

+8.35 pp

Table 15.7. DIAL-URL–2 active discrimination. Rates are macro averages over four true behavior types, with 500 replications in every type–policy cell. The passive schedule is balanced, not random. Active selection increases exact admission at matched probe budget without increasing wrong admission, so the comparison concerns experiment choice rather than additional evidence.

All five preregistered endpoints passed. At eight probes, active selection improved exact admission by 23.35 percentage points while keeping incorrect admission at 1%. At twenty probes it reached 99.65% exact admission, with no incorrect admissions and only 0.35% abstention. MDP false accommodation never exceeded 0.8%, and no composite world was collapsed to an MDP at the largest budget. Median first crossing of the admission threshold fell from eight probes under passive collection to six under active collection.

The intervention trace explains the gain. The learner initially favored the joint probe, which can expose either constructor. As the posterior concentrated, it shifted toward the specialized probe that best separated the remaining alternatives. In the composite world at twenty probes, the mean allocation was 10.44 observation probes, 7.11 duration probes, and 2.45 joint probes; no budget was spent on the weak neutral intervention. The passive control necessarily spent five probes on each.

Experiment: DIAL-URL–2: Active Coalgebraic Discrimination. Input: a posterior over registered behavior functors and a typed intervention menu.
Selection: maximize expected information gain against a balanced passive schedule.
Admission: posterior probability at least 0.95, otherwise abstain.
Outcome: at eight probes, 75.45% active exact admission versus 52.10% passive; at twenty, 99.65% exact with no wrong admissions.
Epistemic status: all preregistered endpoints passed; evidence for active exploratory discrimination, not construction of a theory extension.
Boundary: the grammar, intervention menu, and likelihood are supplied; their operators and new functors are not learned.