lin-0183

14.7 From anonymous anchors to agent infrastructure

A second controlled benchmark replaces schematic anchors with task contracts, plan graphs, tool manifests, router policies, observation schemas, evidence ledgers, citation policies, validators, failure triage, and handoff memory. Six scenarios cover retrieval, extraction, code patching, research synthesis, long-horizon agency, and incident triage.

Exhaustive validation uses 133 calls and scores \(1.000\). LASKO uses 132 cheap bracket probes and 17 validation calls and also scores \(1.000\); random selection uses the same validation budget but averages \(0.322\). The difference is structural localization: read–write dependencies direct validation toward plausible workflow interactions.