Definition. Reward-model validation law [ftip-002U]
Definition. Reward-model validation law [ftip-002U]
A reward-model validation law \(\mathcal D_{\mathrm {rm}}^{\mathrm {val}}\) is a held-out law on comparison observations used to evaluate the fitted score model rather than to update \(\phi \). A declared validation statistic may be Bradley--Terry log loss or the probability that \(r_\phi (x,y^+)>r_\phi (x,y^-)\).
The validation population and acquisition procedure are part of the quantity. Accuracy on comparisons drawn from the same collection process does not establish calibration under a new judge population or agreement with an independent task-success criterion.