Theorem. Finite exponential tilting preserves support [ftip-007C]
AGENTDRAFTED
Fix a prompt \(x\), a finite response set \(\mathcal Y\), a reference policy
\(\pi _{\mathrm {ref}}\), a finite real reward \(r(x,y)\), and \(\beta >0\). Let
\(\pi _r\) be the normalized exponential tilt of Definition [ftip-003O]. Then, for
every \(y\in \mathcal Y\),
\[
\pi _r(y\mid x)>0
\quad \Longleftrightarrow \quad
\pi _{\mathrm {ref}}(y\mid x)>0.
\]
The source policy form appears as equation (4) in
[rafailov2023direct, Section 4]; the displayed result is the finite
corollary derived here. Infinite rewards,
non-normalizable response spaces, approximate optimization, and changes to the
generation mechanism lie outside the statement.