Remark. How the divergence enters alignment objectives [ftip-006Y]

The KL-regularized reward objective appears as equation (3) in Section 3 of [rafailov2023direct]. Behavior-policy ratios, clipped ratios, and KL penalties enter alignment objectives for different purposes.

A coefficient multiplying \(D_{\mathrm {KL}}\) declares an optimization tradeoff. It is neither a guarantee that every sampled ratio is small nor an independent evaluation of the resulting policy.