ch-lincs-rlhf
16 Learning from Structured Preferences
LINCS–RLHF
RLHF usually turns pairwise preferences into a scalar reward and then optimizes that reward. LINCS–RLHF treats the scalar reward as a factorization hypothesis. If comparison log odds circulate around a cycle, no global scalar potential represents the declared preference relation. More reward optimization cannot repair a representation that does not exist.
11. The structural question comes before optimizer choice: should the preference relation be represented by a scalar potential at all? ↩