Definition. Reward--policy reparameterization [rafailov2023direct, Section 4, equation (5)] [ftip-003P]

Rearranging the optimal-policy identity of Definition [ftip-003O] gives

\[ r(x,y) =\beta \log \frac {\pi _r(y\mid x)}{\pi _{\mathrm {ref}}(y\mid x)} +\beta \log Z_r(x). \]

The final term depends on the prompt but not on the response. It cancels from pairwise reward differences, matching the additive non-identifiability of Bradley--Terry scores described in Remark [ftip-002S].