Reward, verifier, and environment failure modes [ftip-00DX]
✍️sourceAGENTDRAFTED
Reward, verifier, and environment failure modes [ftip-00DX]
✍️sourceAGENTDRAFTED
A reward channel, a verifier, and the environment transition law describe different aspects of an interaction. A high score in one channel is not by itself a utility, support, or capability-acquisition statement.
The catalog of [dharna2026aifindsway] supplies heterogeneous empirical counterexamples and provenance, not rates or a general theorem.
1. Reward semantics and environment coupling [ftip-00DY]AGENTDRAFTED
1. Reward semantics and environment coupling [ftip-00DY]AGENTDRAFTED
The same response can be scored by several channels while the environment can assign different transitions or hidden consequences. A score, a state transition, and a hidden consequence are distinct functions of the interaction.
Definition 1.1. Named reward channel [ftip-00DZ]AGENTDRAFTED
Definition 1.1. Named reward channel [ftip-00DZ]AGENTDRAFTED
For a fixed task \(x\), let \(\mathcal Y_x\) be a finite response set and let a named reward channel be a map \(r:\mathcal Y_x\to \mathbb R\). A utility map \(u:\mathcal Y_x\to \mathbb R\) is a separate evaluation quantity.
If an interaction has state space \(\mathcal S\), action space \(\mathcal A\), and transition kernel \(K(\mathord {\cdot }\mid s,a)\), then \(K\) is a third object: changing \(K\) can change consequences without changing \(r\).
Definition 1.2. Verifier channel and false acceptance [ftip-00E0]AGENTDRAFTED
Definition 1.2. Verifier channel and false acceptance [ftip-00E0]AGENTDRAFTED
A verifier channel is a map \(V:\mathcal Y_x\to \{0,1\}\). Given a validity map \(U:\mathcal Y_x\to \{0,1\}\), a false-accept event is
\[F=\{y\in \mathcal Y_x:V(y)=1\ \text {and}\ U(y)=0\}.\]The event depends on the declared verifier and validity test. It is not identified by a scalar reward unless equivalence with that reward criterion is an explicit assumption.
Remark 1.3. What reward-hacking anecdotes establish [ftip-00E1]AGENTDRAFTED
Remark 1.3. What reward-hacking anecdotes establish [ftip-00E1]AGENTDRAFTED
Sections 4.1--4.2 and 7.2--7.3 of [dharna2026aifindsway] catalogue reward, score, environment, and evaluator-target anecdotes. The source gives 26 curated firsthand anecdotes involving more than 100 researchers; it does not provide a common sampling frame, base rates, or causal effect estimates.
Accordingly, this section uses the paper as an empirical counterexample catalog. The finite statements that follow are proved here and do not claim to summarize all training systems.
Theorem 1.4. Uniform proxy disagreement gives two-epsilon regret [ftip-00E2]AGENTDRAFTED
Theorem 1.4. Uniform proxy disagreement gives two-epsilon regret [ftip-00E2]AGENTDRAFTED
Let \(\mathcal Y_x\) be finite, let \(u,r:\mathcal Y_x\to \mathbb R\), and assume \(|r(y)-u(y)|\le \varepsilon \) for every \(y\), with \(\varepsilon \ge 0\). If \(y^\star \) maximizes \(u\) and \(\widehat y\) maximizes \(r\), then
\[u(y^\star )-u(\widehat y)\le 2\varepsilon .\]Indeed, \(u(y^\star )\le r(y^\star )+\varepsilon \le r(\widehat y)+\varepsilon \le u(\widehat y)+2\varepsilon \). This is a finite same-class decision bound, not an optimization, distribution-shift, or capability theorem.
Example 1.5. Equal nominal score, different utility [ftip-00E3]AGENTDRAFTED
Example 1.5. Equal nominal score, different utility [ftip-00E3]AGENTDRAFTED
Take \(\mathcal Y_x=\{a,b\}\), with \(r(a)=r(b)=1\) but \(u(a)=1\) and \(u(b)=0\). Every reward maximizer is a nominal tie, while only \(a\) is utility optimal. The example shows why a score equality does not establish intent or validity.
2. Verifier false accepts [ftip-00E4]AGENTDRAFTED
2. Verifier false accepts [ftip-00E4]AGENTDRAFTED
False acceptance is an event over candidates and a declared validity test. Repeated auditing can bound its occurrence, but a bound on existence is not automatically a bound on the selected output.
Definition 2.1. False-accept event for a candidate [ftip-00E5]AGENTDRAFTED
Definition 2.1. False-accept event for a candidate [ftip-00E5]AGENTDRAFTED
For candidate index \(i\), let \(I_i\in \{0,1\}\) indicate invalidity and let \(A_i\in \{0,1\}\) indicate verifier acceptance. Define \(F_i=\{I_i=1,A_i=1\}\). Write \(q_i=\Pr (I_i=1)\). If \(q_i>0\), write \(\eta _i=\Pr (A_i=1\mid I_i=1)\); if \(q_i=0\), set \(\eta _i=0\). Then \(\Pr (F_i)=q_i\eta _i\).
No independence or identical-distribution hypothesis is part of this definition.
Lemma 2.2. Union bound for repeated false accepts [ftip-00E6]AGENTDRAFTED
Lemma 2.2. Union bound for repeated false accepts [ftip-00E6]AGENTDRAFTED
For finitely many candidate events from Definition 2.1,
\[\Pr \left (\bigcup _{i=1}^{N}F_i\right )\le \sum _{i=1}^{N}\Pr (F_i) =\sum _{i=1}^{N}q_i\eta _i.\]This is the union bound and requires no independence. If every marginal is at most \(\eta \), the right side is at most \(N\eta \).
Example 2.3. Existence of a false accept is not selected-output failure [ftip-00E7]AGENTDRAFTED
Example 2.3. Existence of a false accept is not selected-output failure [ftip-00E7]AGENTDRAFTED
Suppose two candidates are produced: candidate 1 is invalid and falsely accepted, while candidate 2 is valid and rejected by a separate score tie-break. Then \(\bigcup _iF_i\) occurs, but a selection rule that always chooses candidate 2 outputs a valid response. Thus an existence probability and a selected-output probability coincide only after the selection law is specified.
3. Environment and oversight targets [ftip-00E8]AGENTDRAFTED
3. Environment and oversight targets [ftip-00E8]AGENTDRAFTED
An evaluator can be an optimization target while hidden environment consequences remain outside its score. A policy can therefore improve the measured score while changing an unmeasured environmental outcome.
Definition 3.1. Environment exploit [ftip-00E9]AGENTDRAFTED
Definition 3.1. Environment exploit [ftip-00E9]AGENTDRAFTED
Given transition kernel \(K\), reward \(r\), and validity predicate \(U\), an environment exploit is a candidate \(y\in \mathcal Y_x\) whose induced interaction under \(K\) receives high \(r(y)\) while failing \(U\). The definition is relative to the declared kernel, horizon, and validity test; it is not a universal property of a model.
Example 3.2. Same reward, different transition semantics [ftip-00EA]AGENTDRAFTED
Example 3.2. Same reward, different transition semantics [ftip-00EA]AGENTDRAFTED
Let one response receive \(r(y)=1\) under two environments, and let \(W:\mathcal S\to \{0,1\}\) test terminal-state safety. In \(K_1\), its next state \(s_1\) has \(W(s_1)=1\); in \(K_2\), the same observed response enters \(s_2\) with \(W(s_2)=0\). The reward channel alone cannot distinguish the two transition semantics, so equal reward does not certify safe consequences.
Example 3.3. Evaluator-target behavior [ftip-00EB]AGENTDRAFTED
Example 3.3. Evaluator-target behavior [ftip-00EB]AGENTDRAFTED
Let \(V(y)=1\) for every output that contains a visible marker, while \(U\) checks a hidden task condition that the marker does not affect. A policy that optimizes \(V\) can improve its observed score while leaving \(U\) unchanged or worse. This is a finite illustration of evaluator targeting, not a claim about the frequency of such behavior.
Remark 3.4. No base rates or causal estimates [ftip-00EC]AGENTDRAFTED
Remark 3.4. No base rates or causal estimates [ftip-00EC]AGENTDRAFTED
The source catalog does not identify the base rate of exploits, the causal effect of a reward intervention, or a universal relationship between score and utility. Heterogeneous anecdotes therefore motivate audit questions and failure tests, not population-level probabilities.
4. Finite remediation protocol [ftip-00ED]AGENTDRAFTED
4. Finite remediation protocol [ftip-00ED]AGENTDRAFTED
A remediation protocol records what was scored, what was audited, and which environment and verifier versions were used. It may reject candidates before a commit, but its guarantee is only the finite event bound proved below.
Definition 4.1. Audit record [ftip-00EE]AGENTDRAFTED
Definition 4.1. Audit record [ftip-00EE]AGENTDRAFTED
An audit record is a tuple containing candidate identity, evaluator and verifier version, environment or transition-kernel version, invalidity tests, random seed, evidence pointers, and adjudicator decision. A record is complete only when these fields are bound to the candidate and the protocol run.
proposition 4.2. Finite audit-gate bound [ftip-00EF]AGENTDRAFTED
proposition 4.2. Finite audit-gate bound [ftip-00EF]AGENTDRAFTED
Let \(N\in \mathbb N_{\geq 1}\) and \(0\leq \eta \leq 1\). Assume these \(N\) candidates are audited and each invalid candidate is falsely accepted with marginal probability at most \(\eta \). Then the probability that some invalid candidate is accepted is at most \(N\eta \), by the union bound of Lemma 2.2. This conclusion does not require independence.
The statement bounds existence of an accepted invalid candidate. A selected output guarantee additionally requires a typed selection rule and its relation to the audit decisions.
Remark 4.3. Verifier sensitivity and unmeasured failures [ftip-00EG]AGENTDRAFTED
Remark 4.3. Verifier sensitivity and unmeasured failures [ftip-00EG]AGENTDRAFTED
These finite channel distinctions do not prove reward hacking rates, capability acquisition, or safety of a deployed agent. The next questions are to measure verifier sensitivity, environment changes, selection rules, and independent audit power under a declared protocol.