Example. A Direct Preference Optimization log-ratio calculation [ftip-003S]
Example. A Direct Preference Optimization log-ratio calculation [ftip-003S]
Four declared policy probabilities determine the log-ratio margin in one DPO loss term.
Let \(\pi _\theta (y^+\mid x)=0.6\), \(\pi _{\rm ref}(y^+\mid x)=0.3\), \(\pi _\theta (y^-\mid x)=0.2\), and \(\pi _{\rm ref}(y^-\mid x)=0.4\). With \(\beta =1/2\), the DPO logit is \[ z=\beta \left [\log 2-\log (1/2)\right ] =\tfrac 12\log 4=\log 2. \] Hence \(\sigma (z)=2/3\) and the one-pair negative log-likelihood is \(-\log \sigma (z)=\log (3/2)\approx 0.405\).
Substitution into Equation (7) of [rafailov2023direct, Section 4] verifies one loss term. Consistency of the pairwise data, suitability of the reference policy, and improvement in independent utility are separate questions.