Example. A two-response alignment identity [ftip-008S]

Take \(\mathcal Y=\{a,b\}\), \(p=(1/2,1/2)\), \(\beta =1\), and \(r(a)=0\), \(r(b)=\log 3\). Exponential weighting changes the reference law as follows.

The reward gain is

\[ G_p(r;r) =\left (\frac 34-\frac 12\right )\log 3 =\frac 14\log 3. \]

Direct calculation gives

\[ \begin {aligned} D_{\mathrm {KL}}(q_r\Vert p) &=\frac 14\log \frac 12+\frac 34\log \frac 32,\\ D_{\mathrm {KL}}(p\Vert q_r) &=\frac 12\log 2+\frac 12\log \frac 23, \end {aligned} \]

whose sum is \(\frac 14\log 3\). Thus the example checks Theorem [ftip-008O] exactly; it is not an empirical alignment result.