Definition. Direct Preference Optimization loss [rafailov2023direct, Section 4, equation (7)] [ftip-003R]

For comparison observations from \(\mathcal D_{\mathrm {pref}}\), the Direct Preference Optimization loss is

\[ L_{\mathrm {DPO}}(\theta ) =-\mathbb E_{(x,y^+,y^-,j)\sim \mathcal D_{\mathrm {pref}}} \left [ \log \sigma \left (\beta \, \Delta _\theta (x,y^+,y^-) \right ) \right ]. \]

The empirical loss is evaluated on a fixed comparison dataset. The temperature, reference policy, record weights, and handling of ties or abstentions are part of the training specification.