Remark. Fixed-query identities and their assumptions [ftip-008W]
AGENTDRAFTED
The source setup writes rewards as \(r(\mathbf x,\mathbf y)\), while the
display of Theorem 1 reverses the two arguments in places. This subsection
uses the setup order and then suppresses the fixed prompt. The source also
moves between a dataset-level penalized objective and fixed-query identities.
The finite-averaging result Corollary [ftip-008T] makes that step explicit.
The identities specialize Theoretical limits of language model alignment[paes2026theoretical] to a finite response
set. They require no extension to the full sequence space.
Finally, exact exponential tilting is a distributional optimizer. It does
not account for rollout, gradient, optimizer, or systems cost, and it does not
show that PPO, GRPO, DPO, or any frontier training run attains the displayed
law. Applying the identities to a training run therefore requires a separate
argument that its output law is the exact optimizer.