Remark. Alignment criteria across different training routes [ftip-006J]

Ouyang et al. combine demonstrations, comparison-trained rewards, and PPO in [ouyang2022training, Sections 3.1--3.5]. Direct Preference Optimization replaces the learned-reward PPO stage with a preference loss in [rafailov2023direct, Section 4]. The broader alignment-training description is a proposed way to compare such routes; it does not identify their observations or objectives with one another.

The declared behavioral criterion is part of the claim, while the independent evaluation remains a separate object. Their agreement is an empirical or theoretical claim, not a consequence of the word alignment.