tab-dial-url1

15.13.8 Finite-trajectory accommodation

15.13.8 Finite-trajectory accommodation

DIAL-URL–1 replaces exact operators by episodic records containing only observations, actions, rewards, holding times, and next observations. Hidden state, observation maps, and ground-truth types remain evaluator-only. A base categorical predictor estimates reward and next observation from the current observation–action pair. A second predictor also receives the preceding observation, action, and reward. Held-out log-score improvement, measured across episodes, supplies a confidence interval for the observation constructor. A Wilson interval for the frequency of non-unit holding times supplies the duration decision. If either interval intersects its registered practical threshold, the type-level system abstains.

Fifty independently seeded replications were run for each world at 12, 40, and 120 episodes of length 25. Table 15.13.8 reports exact recovery and abstention. Abstention is not counted as correct.

World

Small: exact / abstain

Medium: exact / abstain

Large: exact / abstain

MDP, direct observation

0% / 100%

92% / 8%

98% / 2%

MDP, redundant hidden alias

0% / 100%

88% / 12%

100% / 0%

POMDP, significant alias

0% / 100%

78% / 22%

100% / 0%

SMDP, non-unit duration

90% / 10%

96% / 4%

98% / 2%

Partially observable SMDP

40% / 58%

86% / 14%

100% / 0%

Table 15.6. DIAL-URL–1 finite-trajectory type recovery. Each cell contains 50 replications. Small, medium, and large budgets contain 12, 40, and 120 episodes, respectively. Exact recovery rises with evidence, while abstention at the smallest budgets prevents unsupported accommodation from being counted as discovery.

All preregistered endpoints passed. False accommodation was zero on both MDP controls at every budget. At the large budget every world reached 98–100% exact recovery, and the composite world was never collapsed to an MDP. The mean large-budget history improvement was approximately \(-0.019\) and \(-0.003\) nats per transition in the direct and quotient MDP controls, but \(0.089\) in both partially observable worlds. Thus simulator-level aliases were not themselves treated as evidence for accommodation; only aliases that improved held-out prediction from history triggered the observation constructor.

The small-budget result reveals the intended conservative behavior and one remaining error. Both unit-time MDP controls and the POMDP abstained in every replication because the absence of a rare duration effect was not yet certified. The composite world produced 29 abstentions, 20 correct composite decisions, and one under-accommodation to SMDP when history dependence was incorrectly rejected. This error is preserved rather than absorbed into the abstention count.


Experiment: DIAL-URL–1: Finite-Trajectory Accommodation. Input: observation–action–reward–duration trajectories; no hidden states, operators, or type labels.
Observer: held-out history gain for Obs; Wilson uncertainty for Time.
Admission: minimal registered type or abstention whenever a constructor decision is unresolved.
Outcome: all primary endpoints passed; large-budget recovery was 98–100% with zero false accommodation on MDP controls.
Epistemic status: experiment-level pass; finite-data selection among registered model patterns, not functor invention.
Boundary: the four model patterns, diagnostic interfaces, and thresholds were registered. The experiment does not synthesize a new functor, identify a unique hidden realization, or demonstrate improved control.