Theorem. Post-training cannot distinguish observationally equivalent worlds [ftip-007S]
Theorem. Post-training cannot distinguish observationally equivalent worlds [ftip-007S]
For a finite random variable \(X\), write \(\operatorname {Law}(X)\) for its probability law. If \(h\) maps the value space of \(X\) to a finite set \(\mathcal Y\), its pushforward law \(h_{\#}\operatorname {Law}(X)\) is characterized by \[ \bigl (h_{\#}\operatorname {Law}(X)\bigr )(\{y\}) =\Pr (h(X)=y),\qquad y\in \mathcal Y. \]
Let \(w_0\equiv _{\mathsf P}w_1\). For any finite output set \(\mathcal Y\) and any readout \(h:\mathcal T_{\mathsf P}\to \mathcal Y\),
\[ h_{\#}\operatorname {Law}(T_{w_0}^{\mathsf P}) =h_{\#}\operatorname {Law}(T_{w_1}^{\mathsf P}). \]Thus a final artifact stamp, update decision, or test chosen solely from the declared protocol's transcript has the same distribution in both worlds.
This finite statement follows from the displayed hypotheses.