Remark. Exact optimization on a finite response set [ftip-008G]

Theorems 1 and 2 of the v1 preprint [paes2026theoretical] concern an exactly optimal KL-regularized policy and a fixed query. On a finite response set, their expectations and normalizers are ordinary finite sums.

These exact identities do not establish a comparison among best-of-\(N\), PPO, and GRPO, or a general guarantee for proxy rewards and reward ensembles. Those questions require assumptions beyond exact optimization of one reward.