ora-0081

5.2 System identification and predictive state representations

Fixed doctrine.

System identification supplies the cleanest precursor to UOCL. A learner observes inputs and outputs of an unknown dynamical system and constructs a model sufficient to predict its future behavior. The hypothesis category may contain state-space models, stochastic automata, or controlled processes; morphisms express coordinate changes, simulations, or observational quotients. The presentation is a growing trajectory.

The declaration already assumes that inputs and outputs are typed, temporal composition is meaningful, the dynamics belong to the admitted model class, and the chosen observations can expose the distinctions demanded by the queries. Linear identification may additionally assume finite dimension, controllability, observability, or a noise family. These are not conclusions of the estimator. They define the doctrine in which estimation takes place.

Predictive state representations sharpen the analogy. Rather than postulate a hidden state with privileged ontology, a PSR represents state by predictions of future, action-conditional observation tests [ Littman et al. , 2001 ] . Let a history be \(h=a_1o_1\cdots a_to_t\), and let a test \(\tau =a'_1o'_1\cdots a'_ko'_k\) ask for the probability of receiving the specified observations after issuing the specified actions. For a family of core tests \(\mathcal T=\{ \tau _i\} _{i\in I}\), the predictive state is

\begin{equation} s_{\mathcal T}(h) =\bigl(\Pr (\tau _i\mid h)\bigr)_{i\in I}. \end{equation}
5.1

An action–observation pair extends the presentation from \(h\) to \(hao\), and conditioning updates \(s_{\mathcal T}(h)\) to \(s_{\mathcal T}(hao)\). Thus a PSR has almost exactly the interface of Definition 5.1: controlled stochastic systems form the hypothesis category, histories form the presentation, future tests form the query doctrine, and conditioning supplies the non-anticipatory update.

The induced observational quotient is especially important. Two histories are predictively equivalent when

\begin{equation} h\sim _{\mathcal T}h' \quad \Longleftrightarrow \quad \Pr (\tau \mid h)=\Pr (\tau \mid h') \quad \text{for every }\tau \in \mathcal T. \end{equation}
5.2

The learner identifies the quotient visible through \(\mathcal T\), not a unique hidden-state ontology. Enlarging the test family refines the quotient; finding a finite sufficient family amounts to finding a compact presentation of the relevant predictive world.

This test-based view has a deterministic precursor in the diversity representation of Rivest and Schapire. Their learner represents a finite automaton by equivalence classes of experiments rather than by enumerating its global states [ Rivest and Schapire , 1994 ] . A PSR may therefore be read as a stochastic and, in common finite-dimensional realizations, linear analogue: Boolean outcomes of deterministic tests are replaced by conditional probabilities of action–observation tests. The analogy concerns the predictive interface; it does not identify the two learning algorithms or their assumptions.

Proposition 5.5 PSRs are query-relative UOCL

Fix a class of controlled stochastic processes and a test family \(\mathcal T\) closed under the one-step extensions required by the update. If the PSR update is well defined on equivalence classes \([h]_{\mathcal T}\), then PSR identification is a UOCL specialization whose identified object is the action of one-step extensions on the predictive quotient of histories.

Proof

The hypothesis class, presentation, and query doctrine have just been specified. Equation 5.2 supplies observational equivalence. Closure under the required extensions makes conditioning a well-defined update on equivalence classes, while predictive accuracy or limiting identification supplies the success criterion. The specialization criterion, Proposition 5.3, then applies.

This also marks the boundary between UOCL and UODL. The actions occurring in a PSR test are controlled input symbols used to interrogate dynamics. They do not by themselves say which action should be chosen. Preferences, consequences, feasibility, and a comparison observer must still be added to obtain a decision declaration. Conversely, a query family sufficient for one-step prediction may fail to preserve what is needed for long-horizon control.

Doctrinal boundary.

A failed parameter fit is ordinarily repaired inside the system doctrine. A persistent failure of every admitted realization may instead challenge stationarity, the state dimension, the observation interface, the noise model, or the existence of a finite sufficient test family. PSR learning becomes doctrine-learning only if such alternatives are themselves typed and available to the update mechanism.