Remark. Assumptions behind the DPO derivation [rafailov2023direct, Section 4 and Appendix A.1--A.2] [ftip-003T]

The derivation uses a fixed reference policy, a finite KL-regularized optimum, and a Bradley--Terry model for pairwise preferences. The relevant policy probabilities must be positive wherever their logarithmic ratios are evaluated.

The algebra eliminates the prompt-dependent reward offset, not every source of reward-model misspecification. Heterogeneous or nontransitive preferences, adaptive data collection, support mismatch, and finite-sample optimization remain separate questions.