Definition. Policy-ratio preference probability [rafailov2023direct, Section 4, equation (6)] [ftip-003Q]

Substituting the reward--policy reparameterization into the Bradley--Terry law yields

\[ \Pr _\theta (y^+\succ y^-\mid x) =\sigma \left (\beta \, \Delta _\theta (x,y^+,y^-) \right ), \]

where \(\Delta _\theta \) is defined in Notation [ftip-003M]. The prompt-dependent normalizer cancels because the two responses share the same prompt.