Remark. Reward-model validation depends on its population [ftip-006O]

Ouyang et al. hold out comparison data and report reward-model validation accuracy and loss in [ouyang2022training, Section 3.5 and Appendix C.2]. Representing the validation population by a probability law makes the estimand depend explicitly on the judge population and acquisition procedure.

This law belongs to reward-model validation. It is not the independent task evaluation used later to compare post-training protocols.