lin-0207

16.7 What the experiments establish

Experiment: LINCS-RLHF controlled preference suite

E1–E4 are controlled preference simulations in which independent random seeds are the replication unit and repeated pairwise labels are sampled from a known reciprocal population matrix. E1 varies labels per edge and cyclic strength to separate null calibration from power. E2 assigns opposite cycles to two annotator strata and compares pooled with localized decisions. E3 uses complete five- and seven-response games and evaluates fitted policies against the population matrix. E4 varies observed graph support among trees, one-cycle graphs, sparse connected graphs, and complete graphs. Oracle population quantities score detection and exploitability but are unavailable to the fitted gate and policies.

Study

Retained result

Failure or boundary

Architectural consequence

E1: calibration and power

null calibration within \(0.015\) for \(n\geq 40\); power at least \(0.930\) in the registered high-information regime

near-zero power in deliberately weak cells

admit only at declared support and resolution

E2: opposite populations

equal-mixture pooling falsely admits \(99.6\% \); localization changes mean exploitability \(0.089\to 0.047\)

four low-information cells worsen by at most \(0.0042\)

preserve annotator strata before aggregation

E3: larger games

mean exploitability: scalar \(0.084\), uniform fallback \(0.094\), relational fallback \(0.026\)

55 of 144 every-cell criteria fail, mostly balanced games

relational fallback is preferred, not universally dominant

E4: incomplete graphs

all trees labeled unidentifiable; false rejection \(0.008\) at \(\alpha =.01\) on cyclic-support graphs

power depends on signal alignment, not cycle rank alone

collect repeated comparisons on informative cycles

Table 16.1 The registered sequence separates calibration, localization, fallback quality, and graph identifiability. Negative cells remain visible.

A public-data audit sharpens the collection requirement. OpenAI summarization-feedback batch 3 contained prompt-level cycles after support thresholding, but no worker-by-prompt stratum retained a cycle once each edge required even two labels. It can support a prompt-level feasibility analysis, not the annotator-local mechanism tested in E2. An audit-ready preference dataset must preserve stable response identities, repeated connected cycles, and annotator metadata.

Admission contract

Scalar policy optimization is admitted only when the declared comparison graph contains estimable cycle witnesses, the stratum-level test does not reject at its frozen threshold, support covers the intended deployment graph, and multiplicity is handled at the level of the scientific claim. Otherwise the controller collects data, weakens the target, abstains, or retains the relational game.