Definition. Policy-ratio preference probability [rafailov2023direct, Section 4, equation (6)] [ftip-003Q]
Definition. Policy-ratio preference probability [rafailov2023direct, Section 4, equation (6)] [ftip-003Q]
Substituting the reward--policy reparameterization into the Bradley--Terry law yields
\[ \Pr _\theta (y^+\succ y^-\mid x) =\sigma \left (\beta \, \Delta _\theta (x,y^+,y^-) \right ), \]where \(\Delta _\theta \) is defined in Notation [ftip-003M]. The prompt-dependent normalizer cancels because the two responses share the same prompt.