ch-lincs-rlhf

16 Learning from Structured Preferences

LINCS–RLHF

RLHF usually turns pairwise preferences into a scalar reward and then optimizes that reward. LINCS–RLHF treats the scalar reward as a factorization hypothesis. If comparison log odds circulate around a cycle, no global scalar potential represents the declared preference relation. More reward optimization cannot repair a representation that does not exist.

11. The structural question comes before optimizer choice: should the preference relation be represented by a scalar potential at all?