Remark. Reward modeling, alignment criteria, and policy optimization [ftip-003B]

The RLHF pipeline makes three logically separate choices. The comparison data specify which judgments were observed; the reward model specifies how those observations are represented and generalized; the policy optimizer specifies how sampled actions change the model.

An alignment criterion lies outside this chain unless it is identified with the training reward by assumption. A reward model can predict its held-out comparisons while failing under policy-induced distribution shift, and PPO can increase the learned reward while independent utility stays fixed or falls.