lin-0209

Further reading

The standard scalar pipeline combines preference-based reinforcement learning [ Christiano et al. , 2017 ] with Bradley–Terry comparison models [ Bradley and Terry , 1952 ] ; direct preference optimization retains the same reward-difference form inside a policy objective [ Rafailov et al. , 2023 ] . Reward overoptimization [ Gao et al. , 2023 ] concerns what happens when a learned scalar proxy is optimized too strongly. LINCS–RLHF asks the logically prior question of whether the comparison field is representable by any scalar reward on the declared support.

Graph-Hodge ranking [ Jiang et al. , 2011 ] supplies the gradient/cycle decomposition. Relational alternatives include minimax RLHF [ Swamy et al. , 2024 ] and general preference models [ Zhang et al. , 2025b ] . Recent hybrid work explicitly separates transitive and cyclic components and optimizes the resulting game dynamically [ Huang et al. , 2026 ] . Together these sources suggest a growing literature in which scalar reward, cyclic structure, population heterogeneity, and game solutions are treated as distinct modeling choices rather than forced into one ranking.