ora-0008

0.5 Embodied action as a discovered world

Learning to reach, grasp, stand, and walk presents a different kind of fragmentation. Here an action changes the evidence that will arrive next. A turn of the head changes the visual field; a reach changes both proprioception and the location of an object; a failed step changes balance and the set of available recoveries. The learner is simultaneously discovering a world and perturbing its presentation.

Spelke’s systems for objects, form, and navigable space supply plausible initial observers. They distinguish a bounded persisting thing from a passing texture, a support relation from coincidence, and a stable layout from the learner’s changing viewpoint. Those distinctions do not contain the skill of grasping a particular cup. They make the relevant regularities available for learning.

Embodied composition can be seen at three scales:

  1. sensorimotor fragments, such as orienting the hand or shifting weight;

  2. reusable skills, such as reaching, grasping, releasing, or recovering balance; and

  3. activities assembled from skills, such as picking up a cup and carrying it without spilling.

The crucial word is when. Grasping composes with lifting only when the grasp is secure; stepping composes with transferring weight only while a support condition holds. A compositional action world must therefore encode domains of admissibility and not just concatenate motor commands.

Learned world models in machine learning separate, at least conceptually, a predictive model of dynamics from the controller that uses it [ Ha and Schmidhuber , 2018 ] . ORACLE sharpens that separation. Learning that an action leads to an observation is not yet learning which actions compose, which histories are equivalent, or which failure calls for a local repair. Assimilation refines a skill inside a stable organization; accommodation may split one skill into context-dependent variants, introduce a new intermediate state, or withdraw an unsafe composite.