Theorem. Cross-reward gain is a base-law covariance [ftip-008Q]

For a finite KL-alignment instance,

\[ G_p(r;s) =\operatorname {Cov}_p\left (s,\frac {a_r}{Z_r}\right ). \]

This is the finite form of Theorem 2, equation (3.4), in Theoretical limits of language model alignment[paes2026theoretical]. Its proof is in Appendix B.1, equations (B.5)--(B.7).