Feedback identifiability and fixed records [ftip-007P]
✍️sourceAGENTDRAFTED
Feedback identifiability and fixed records [ftip-007P]
✍️sourceAGENTDRAFTED
A training protocol can respond only to distinctions present in its observations. The following finite setup makes that restriction precise before considering stronger, model-specific limits on preference data.
Definition 1. Finite declared feedback protocol [ftip-007Q]AGENTDRAFTED
Definition 1. Finite declared feedback protocol [ftip-007Q]AGENTDRAFTED
Fix a horizon \(N\geq 0\), finite action sets \(\mathcal A_n\), finite feedback sets \(\mathcal F_n\), and a finite internal-seed set \(\mathcal U\) with a declared law \(\lambda \). A finite declared protocol \(\mathsf P\) consists of selection maps
\[ a_n:\mathcal U\times \prod _{j<n}(\mathcal A_j\times \mathcal F_j) \longrightarrow \mathcal A_n \qquad (0\leq n<N). \]An action may specify the task, rollout record, feedback request, or sampling choice made at that round. The protocol can depend on earlier exposed feedback and its declared seed, but it cannot depend on a latent world field that has not entered those arguments.
Definition 2. Finite feedback world [ftip-008C]AGENTDRAFTED
Definition 2. Finite feedback world [ftip-008C]AGENTDRAFTED
For the action and feedback sets of Definition 1, a finite feedback world \(w\) supplies, for each round \(0\leq n<N\), a probability kernel \(K_{n,w}\) from a generic action--feedback history \((a_0,f_0,\ldots ,a_n)\) ending in \(a_n\in \mathcal A_n\) to \(\mathcal F_n\). Thus, whenever \(a_j\in \mathcal A_j\) and \(f_j\in \mathcal F_j\),
\[ K_{n,w}(\,\cdot \mid a_0,f_0,\ldots ,a_n) \in \Delta (\mathcal F_n). \]Two worlds may agree on every feedback kernel queried by a protocol while differing in latent utility, unrequested labels, or unobserved environment facts. Those latent fields are not feedback until a declared query exposes them.
Definition 3. Observed feedback transcript [ftip-008D]AGENTDRAFTED
Definition 3. Observed feedback transcript [ftip-008D]AGENTDRAFTED
Fix a finite declared protocol \(\mathsf P\) from Definition 1 and a finite feedback world \(w\) from Definition 2. Draw \(U\sim \lambda \) and, for \(0\leq n<N\), set
\[ \begin {aligned} A_n&=a_n(U,A_0,F_0,\ldots ,A_{n-1},F_{n-1}),\\ F_n&\sim K_{n,w}(\,\cdot \mid A_0,F_0,\ldots ,A_n). \end {aligned} \]The resulting observed feedback transcript is
\[ T_w^{\mathsf P}=(U,A_0,F_0,\ldots ,A_{N-1},F_{N-1}). \]All protocol randomness is included in \(U\). Immutable rollout records and typed feedback events may be components of the finite action and feedback sets, as in Definition [ftip-0051] and Definition [ftip-0053].
Definition 4. Observational equivalence for a declared protocol [ftip-007R]AGENTDRAFTED
Definition 4. Observational equivalence for a declared protocol [ftip-007R]AGENTDRAFTED
Let \(\mathcal T_{\mathsf P}\) be the finite set of possible transcripts for a declared protocol \(\mathsf P\). Two feedback worlds \(w_0,w_1\) are observationally equivalent for \(\mathsf P\), written \(w_0\equiv _{\mathsf P}w_1\), when
\[ \Pr (T_{w_0}^{\mathsf P}=t)=\Pr (T_{w_1}^{\mathsf P}=t) \qquad \text {for every }t\in \mathcal T_{\mathsf P}. \]The relation is protocol-relative. Another protocol may issue a different feedback request and thereby separate the same worlds. Equality only on the realized transcript is weaker than this definition, which compares the whole finite transcript law.
Theorem 5. Post-training cannot distinguish observationally equivalent worlds [ftip-007S]AGENTDRAFTED
Theorem 5. Post-training cannot distinguish observationally equivalent worlds [ftip-007S]AGENTDRAFTED
For a finite random variable \(X\), write \(\operatorname {Law}(X)\) for its probability law. If \(h\) maps the value space of \(X\) to a finite set \(\mathcal Y\), its pushforward law \(h_{\#}\operatorname {Law}(X)\) is characterized by \[ \bigl (h_{\#}\operatorname {Law}(X)\bigr )(\{y\}) =\Pr (h(X)=y),\qquad y\in \mathcal Y. \]
Let \(w_0\equiv _{\mathsf P}w_1\). For any finite output set \(\mathcal Y\) and any readout \(h:\mathcal T_{\mathsf P}\to \mathcal Y\),
\[ h_{\#}\operatorname {Law}(T_{w_0}^{\mathsf P}) =h_{\#}\operatorname {Law}(T_{w_1}^{\mathsf P}). \]Thus a final artifact stamp, update decision, or test chosen solely from the declared protocol's transcript has the same distribution in both worlds.
Proof.
Proof.
For every \(y\in \mathcal Y\), finiteness gives
\[ \begin {aligned} \Pr \left (h(T_{w_0}^{\mathsf P})=y\right ) &=\sum _{\substack {t\in \mathcal T_{\mathsf P}\\h(t)=y}} \Pr (T_{w_0}^{\mathsf P}=t)\\ &=\sum _{\substack {t\in \mathcal T_{\mathsf P}\\h(t)=y}} \Pr (T_{w_1}^{\mathsf P}=t) =\Pr \left (h(T_{w_1}^{\mathsf P})=y\right ). \end {aligned} \]The middle equality is observational equivalence from Definition 4. These point probabilities determine the two pushforward laws.
This finite statement follows from the displayed hypotheses.
Remark 6. A no-free-feedback result, not a no-learning result [ftip-007T]AGENTDRAFTED
Remark 6. A no-free-feedback result, not a no-learning result [ftip-007T]AGENTDRAFTED
Theorem 5 says that the declared observations supply no information that distinguishes \(w_0\) from \(w_1\). It does not say that the output must equal the initial model, that the parameters cannot change, or that performance cannot improve in both worlds. Pretraining, inductive bias, computation on the observed records, and generalization may still produce an improvement shared by the two worlds.
Calling the theorem a no-learning result would therefore erase the central condition: only world-dependent conclusions unavailable from the common transcript are ruled out.
Corollary 7. Replaying a fixed pool supplies no new feedback information [ftip-007U]AGENTDRAFTED
Corollary 7. Replaying a fixed pool supplies no new feedback information [ftip-007U]AGENTDRAFTED
Fix a realized finite replay pool \(d\) as in Definition [ftip-0056]. Let \(\mathcal V\) be a finite seed set, and let \(V\in \mathcal V\) be a replay-and-update seed with the same declared law in two feedback worlds and whose law does not depend on the world. If a replay-only procedure makes no new environment or feedback query, then its output has the form \(Y=h(d,V)\). The law of \(Y\) is the same in the two worlds.
Proof.
Proof.
For every output \(y\),
\[ \Pr (h(d,V)=y)= \sum _{\substack {v\in \mathcal V\\h(d,v)=y}}\Pr (V=v). \]The fixed pool, seed law, and readout are identical in the two worlds, so the displayed sum is identical. Equivalently, this is the pushforward argument of Theorem 5 applied after conditioning on the realized pool.
This finite statement follows from the displayed hypotheses.
Remark 8. Reuse can change optimization without enlarging evidence [ftip-007V]AGENTDRAFTED
Remark 8. Reuse can change optimization without enlarging evidence [ftip-007V]AGENTDRAFTED
A replay schedule can reweight records, reduce optimization error on the fixed pool, or produce a different update proposal in the sense of Definition [ftip-006C]. Those are genuine computational effects. They do not add a label, preference, verifier result, or environment transition to the fixed evidence. Fresh optimizer randomness is likewise not feedback about the latent world.
This distinction prevents sample reuse from being counted as feedback acquisition. It does not imply that replay is useless; it isolates the source of any benefit as further computation on already acquired records.
Example 9. Counterexample: An off-query utility reversal [ftip-007W]AGENTDRAFTED
Example 9. Counterexample: An off-query utility reversal [ftip-007W]AGENTDRAFTED
Let the response set be \(\{a,b\}\). A one-round protocol queries only \(x_0\); both worlds return the preference \(a\succ b\). The protocol therefore has the same transcript law in both worlds. Fix a deterministic transcript-to-policy readout that chooses \(a\) also at an unqueried \(x_1\). The worlds agree on every observed field but reverse utility at \(x_1\):
The transcript laws are identical, so Theorem 5 applies, while the deterministic readout returns the same output in both worlds. That output is optimal in \(w_+\) and suboptimal in \(w_-\) on \(x_1\). The example does not show that generalization always fails. It shows that a claim about unqueried utility requires an assumption connecting observed feedback to that utility.
Remark 10. From finite indistinguishability to preference-data limits [ftip-007X]AGENTDRAFTED
Remark 10. From finite indistinguishability to preference-data limits [ftip-007X]AGENTDRAFTED
[zhao2025limits, Theorems 3.3--3.5] study a more structured post-training model. They give ordinal-preference distortion lower bounds, including a lower bound under Bradley--Terry noise with linear scores, and a positive result using a limited number of cardinal queries. These results depend on the routing model, utility class, and query budget developed in § [ftip-008X].
Theorem 5 gives a finite pushforward theorem, while Example 9 gives a counterexample to identification from an incomplete transcript. They are neither proofs nor special cases of the paper's distortion results.