lin-0231

19.4 E\(_0\): deterministic structural perturbations

The E\(_0\) study asks whether paired SCoLT profiles distinguish declared null changes from material changes to corpus-derived tickets. The evaluation uses all tickets produced from the 340-document Web-discourse corpus of Habernal and Gurevych and the 112 English argumentative microtexts of Peldszus and Stede, with no fitted parameters and therefore no train/test split.

Null controls changed formatting or reversed independent grounds. Material probes dropped a claim, warrant, ground, or rebuttal; strengthened a qualifier; or swapped a warrant or source binding. Each material probe changed one declared obligation. Inapplicable probes were omitted rather than counted as negatives; swaps used the next non-equivalent populated ticket within the same corpus.

The one-point comparator inspected only the perturbed ticket and detected missing required fields. The paired LINCS profile compared anchor and perturbation after applying the quotient, then emitted typed changes for claim, grounds, warrant, qualifier, rebuttal, and source.

Measure

One-point base gate

SCoLT

Sensitivity to material change

0.4090

1.0000

Specificity on null controls

0.9783

1.0000

Balanced accuracy

0.6936

1.0000

F1

0.5786

1.0000

Exact obstruction localization

not available

1.0000

Table 19.2 E\(_0\) results on 564 corpus-derived tickets and 4,283 perturbation pairs.

The observed balanced-accuracy improvement was \(0.3064\), exceeding the pre-specified \(0.25\) criterion. The strongest base-gate blind spots were populated-but-swapped warrants, changed source bindings, strengthened qualifiers, and removed rebuttals.

Repair controls made the structural distinction sharper. LINCS admitted every exact restoration and rejected every generic nonempty surface patch. The base completeness gate admitted \(96.99\% \) of those superficial patches.

Experiment: what E\(_0\) validates

E\(_0\) validates the executable distinction between checking whether a ticket is populated and checking whether its typed obligations remain invariant under a declared change. It also validates the two quotient-null controls, typed localization, and blockwise rejection of nonempty but unwarranted repairs.

Boundary

The perturbations and evaluator share a typed structured representation. The perfect LINCS result therefore validates factorization, quotient, localization, and admission plumbing. It does not demonstrate robustness to natural paraphrases, noisy extraction, learned repair, answer correctness, automatic-differentiation semantics for language, or preference alignment.