Notation. Translating the alignment source notation [ftip-008H]
Notation. Translating the alignment source notation [ftip-008H]
The source writes \(\lambda >0\) for the KL penalty. Here we write \(\beta >0\), matching the DPO notation in Notation [ftip-003M]--Definition [ftip-003O]. This avoids collision with the generalized-advantage parameter in Definition [ftip-003E] and the protocol law in Definition [ftip-007Q].
The source's gain \(\Delta (r,r')\) is written \(G_p(r;s)\): \(r\) is the reward used to tilt the reference law \(p\), while \(s\) is the reward used to evaluate the tilted law. This avoids collision with both the simplex notation \(\Delta (X)\) of Notation [ftip-000A] and the paired DPO margin of Notation [ftip-003M]. For a finite set \(\mathcal Y\), a law \(p\in \Delta (\mathcal Y)\), and functions \(f,g:\mathcal Y\to \mathbb R\), set
\[ \operatorname {Cov}_p(f,g) =\mathbb E_{y\sim p}[f(y)g(y)] -\mathbb E_{y\sim p}[f(y)]\, \mathbb E_{y\sim p}[g(y)]. \]