lin-0170

13.5 Controlled and language-model evidence

Three metrics recur below. Pair RMSE compares the outputs of \(i\! \to j\) and \(j\! \to i\) on the same held-out inputs. Triple RMSE aggregates disagreement across permutations of three adapters. The reported “slope” is the fitted log–log exponent of order discrepancy against a common adapter-strength multiplier; the linear proposition predicts an exponent of two near zero strength. Task \(\Delta \) is always reported separately so that two equally ineffective adapter orders cannot count as a successful repair.

The controlled behavioral experiment first isolates the mechanism: at \(\lambda _c=0.03\), the bracket falls \(1667\times \), logit-order sensitivity \(1.72\times 10^7\), and style-order sensitivity \(1.89\times 10^6\), while task metrics remain perfect.

Across a six-setting pilot corpus–backbone matrix at the same weight, commutators fall \(14.5\times \)–\(22.1\times \) and output-order sensitivity falls \(27.7\times \)–\(77.8\times \). Five of six next-token losses change by less than \(0.01\).

Setting

Bracket

Pair RMSE

Triple RMSE

Task \(\Delta \)

Slope

PTB / Transformer

\(21.73\times \)

\(7.17\times \)

\(7.31\times \)

\(+0.0005\pm 0.0008\)

\(1.64\)

WikiText-2 / KET incidence

\(35.46\times \)

\(10.94\times \)

\(11.93\times \)

\(+0.0026\pm 0.0003\)

\(1.71\)

Table 13.1 Five-seed theory-driven tests at \(\lambda _c=0.03\). Every paired seed improves; linear operator slopes are approximately \(2.03\), as predicted.

The network-output slopes of \(1.48\)–\(1.71\) are below the exact linear slope near two. This gap is useful evidence: it identifies the effect of nonlinearities and repeated insertion sites rather than pretending that the matrix model is exact end to end.