Example. Same reward, different transition semantics [ftip-00EA]

Let one response receive \(r(y)=1\) under two environments, and let \(W:\mathcal S\to \{0,1\}\) test terminal-state safety. In \(K_1\), its next state \(s_1\) has \(W(s_1)=1\); in \(K_2\), the same observed response enters \(s_2\) with \(W(s_2)=0\). The reward channel alone cannot distinguish the two transition semantics, so equal reward does not certify safe consequences.