Theorem. Exact reward gain equals scaled Jeffreys divergence [ftip-008O]

For a finite KL-alignment instance,

\[ G_p(r;r)=\beta J(p,q_r). \]

This is the finite form of [paes2026theoretical, Theorem 1, equation (3.2), with proof in Appendix B.1].