ifc-0173
13.1 What is creative about reinforcement learning?
Let a conventional decision declaration be
where \(\mathcal X\) and \(\mathcal A\) describe the state and action types, \(P\) is a transition law, \(R\) is a reward interface, \(\gamma _{\mathrm{disc}}\) is a discount convention, and \(\mathcal O\) records the observer available to the agent. A policy \(\pi \) and critic \(Q\) are models inside this declaration. Ordinary learning moves the pair
while keeping \(\mathbb S_{\mathrm{RL}}\) fixed.
That is assimilation in the sense used throughout this book. A proto-accommodation proposal changes the declaration or the language in which its policies are assembled:
Examples include adding a latent regime coordinate, forming a temporally extended option, separating one action type into two, introducing an object relation, or changing which observations are retained as persistent state. The creative delta is the newly expressible decision structure, not merely a higher-return parameter vector. Installation as accommodation occurs only after finite realization and independent admission.
. Intrinsic motivation changes which experiences an agent seeks. RELIC changes which policy structures the agent may construct, but admits such a change only after an independent test of value.
This distinction does not diminish exploration. Novelty, surprise, and information gain may be excellent selectors of experiments. Their semantic role is nevertheless different from a declaration extension. A count bonus can lead an agent toward an unfamiliar cupboard without teaching it that “find, acquire, transport, place” should be retained as one reusable composition.