Theorem. Post-training cannot distinguish observationally equivalent worlds [ftip-007S]
AGENTDRAFTED
For a finite random variable \(X\), write \(\operatorname {Law}(X)\) for its
probability law. If \(h\) maps the value space of \(X\) to a finite set
\(\mathcal Y\), its pushforward law \(h_{\#}\operatorname {Law}(X)\) is
characterized by
\[
\bigl (h_{\#}\operatorname {Law}(X)\bigr )(\{y\})
=\Pr (h(X)=y),\qquad y\in \mathcal Y.
\]
Let \(w_0\equiv _{\mathsf P}w_1\). For any finite output set \(\mathcal Y\)
and any readout \(h:\mathcal T_{\mathsf P}\to \mathcal Y\),
\[
h_{\#}\operatorname {Law}(T_{w_0}^{\mathsf P})
=h_{\#}\operatorname {Law}(T_{w_1}^{\mathsf P}).
\]Thus a final artifact stamp, update decision, or test chosen solely from
the declared protocol's transcript has the same distribution in both worlds.
This finite statement follows from the displayed hypotheses.