Theorem. Finite exponential tilting preserves support [ftip-007C]

Fix a prompt \(x\), a finite response set \(\mathcal Y\), a reference policy \(\pi _{\mathrm {ref}}\), a finite real reward \(r(x,y)\), and \(\beta >0\). Let \(\pi _r\) be the normalized exponential tilt of Definition [ftip-003O]. Then, for every \(y\in \mathcal Y\),

\[ \pi _r(y\mid x)>0 \quad \Longleftrightarrow \quad \pi _{\mathrm {ref}}(y\mid x)>0. \]

The source policy form appears as equation (4) in [rafailov2023direct, Section 4]; the displayed result is the finite corollary derived here. Infinite rewards, non-normalizable response spaces, approximate optimization, and changes to the generation mechanism lie outside the statement.