Definition. Pairwise reward-model loss [ouyang2022training, Section 3.5, equation (1), and Appendix C.2] [ftip-002R]

For comparison observations sampled from \(\mathcal D_{\mathrm {pref}}\), the pairwise reward-model loss is

\[ L_{\mathrm {rm}}(\phi ) =\mathbb E_{(x,y^+,y^-,j)\sim \mathcal D_{\mathrm {pref}}} \left [-\log \sigma \left ( r_\phi (x,y^+)-r_\phi (x,y^-) \right )\right ]. \]

The empirical objective replaces the expectation by a declared weighting of the collected comparisons. It fits score differences under the Bradley--Terry law of Definition [ftip-002Q]; it does not observe an absolute reward target for either response.