Evaluation transport and matched protocols [ftip-00F9]
✍️sourceAGENTDRAFTED
Evaluation transport and matched protocols [ftip-00F9]
✍️sourceAGENTDRAFTED
This section compares post-training artifacts only through declared evaluation laws, inference protocols, utilities, seeds, and budgets. It proves finite transport statements locally and treats long-horizon research-state records as empirical evidence, not as theorems.
1. Evaluation laws and couplings [ftip-00FA]AGENTDRAFTED
1. Evaluation laws and couplings [ftip-00FA]AGENTDRAFTED
Convention 1.1. A named evaluation experiment [ftip-00FB]AGENTDRAFTED
Convention 1.1. A named evaluation experiment [ftip-00FB]AGENTDRAFTED
A named evaluation experiment records the artifact, task law, inference protocol, evaluator, seed kernel, budget, and sampling unit. Two reported scores are comparable only after these coordinates and their intended target have been stated.
Definition 1.2. Evaluation law [ftip-00FC]AGENTDRAFTED
Definition 1.2. Evaluation law [ftip-00FC]AGENTDRAFTED
For a fixed evaluation interface from Convention [ftip-005D], an evaluation law is a probability measure \(\mathsf Q\) on the finite outcome space \(\mathcal Z=\mathcal X_{\mathsf T_{\rm ev}}\times \mathcal O_{\mathsf T_{\rm ev}}\) with random pair \((X,O)\sim \mathsf Q\). Its score for artifact \(M\) is a bounded measurable map \(s_M:\mathcal Z\to [0,1]\). Write \(J_{\mathsf Q}(M)=\mathbb E_{\mathsf Q}[s_M(X,O)]\).
Definition 1.3. Matched evaluation pair [ftip-00FD]AGENTDRAFTED
Definition 1.3. Matched evaluation pair [ftip-00FD]AGENTDRAFTED
A matched evaluation pair is \((\mathsf Q,\mathsf I,b,\Lambda )\) and \((\mathsf Q',\mathsf I',b',\Lambda ')\) together with a coupling \((Z,Z')\) whose marginals are the two laws. The pair is matched on the declared coordinates when the task, utility, evaluator version, and sampling unit are shared; any changed coordinate is named.
Definition 1.4. Paired finite score estimator [ftip-00FE]AGENTDRAFTED
Definition 1.4. Paired finite score estimator [ftip-00FE]AGENTDRAFTED
For iid coupled draws \((Z_j,Z'_j)_{j=1}^m\), define \(\widehat \Delta _m=m^{-1}\sum _{j=1}^m(s_M(Z_j)-s_{M'}(Z'_j))\). The pairing is part of the estimator specification; it does not silently identify the two evaluation laws.
Theorem 1.5. Unbiasedness of the paired score estimator [ftip-00FF]AGENTDRAFTED
Theorem 1.5. Unbiasedness of the paired score estimator [ftip-00FF]AGENTDRAFTED
Under the iid coupling in Definition 1.4, \(\mathbb E[\widehat \Delta _m]=J_{\mathsf Q}(M)-J_{\mathsf Q'}(M')\).
Proof.
Proof.
Linearity of expectation gives the displayed identity.
\[ \mathbb E[\widehat \Delta _m] =m^{-1}\sum _j\left (\mathbb E[s_M(Z_j)]- \mathbb E[s_{M'}(Z'_j)]\right ). \]The coupling fixes the joint sampling but does not change either marginal, so the two terms are the stated expectations.
Example 1.6. Common seeds can reduce paired variance [ftip-00G0]AGENTDRAFTED
Example 1.6. Common seeds can reduce paired variance [ftip-00G0]AGENTDRAFTED
If \(s_M(Z)=s_{M'}(Z')\) almost surely under a common-seed coupling, then \(\widehat \Delta _m=0\) for every sample even when independent sampling would have nonzero variance. This is a variance statement, not evidence that the two artifacts have equal behavior on a different law.
Remark 1.7. A coupling does not merge estimands [ftip-00G1]AGENTDRAFTED
Remark 1.7. A coupling does not merge estimands [ftip-00G1]AGENTDRAFTED
The result in Theorem 1.5 follows from its displayed hypotheses. A shared seed can make a paired comparison precise while the changed task law, evaluator, or inference budget still changes the estimand. No causal effect follows without a declared intervention and matched coordinates.
2. Utility transport [ftip-00G2]AGENTDRAFTED
2. Utility transport [ftip-00G2]AGENTDRAFTED
Definition 2.1. Bounded utility perturbation [ftip-00G3]AGENTDRAFTED
Definition 2.1. Bounded utility perturbation [ftip-00G3]AGENTDRAFTED
On a common finite outcome law \(\mathsf Q\), let each artifact \(M\) have score maps \(s_M,t_M:\mathcal Z\to [0,1]\). They have perturbation radius \(\varepsilon =\sup _{M,\,z\in \operatorname {supp}(\mathsf Q)} |s_M(z)-t_M(z)|\) when \(0\leq \varepsilon \leq 1\); write \(J_s(M)=\mathbb E_{\mathsf Q}[s_M]\) and \(J_t(M)=\mathbb E_{\mathsf Q}[t_M]\).
Theorem 2.2. Direct utility transport bound [ftip-00G4]AGENTDRAFTED
Theorem 2.2. Direct utility transport bound [ftip-00G4]AGENTDRAFTED
Under Definition 2.1, every artifact \(M\) satisfies \(|J_s(M)-J_t(M)|\leq \varepsilon \).
Proof.
Proof.
Pointwise bounds give \(-\varepsilon \leq s_M(z)-t_M(z)\leq \varepsilon \). Taking expectations preserves both inequalities.
Corollary 2.3. Transporting an artifact comparison [ftip-00G5]AGENTDRAFTED
Corollary 2.3. Transporting an artifact comparison [ftip-00G5]AGENTDRAFTED
If two artifacts \(M_1,M_0\) are scored by both \(s\) and \(t\) under the same law and \(\lVert s-t\rVert _{\infty ,\mathsf Q}\leq \varepsilon \), then the gain difference obeys \(|(J_s(M_1)-J_s(M_0))-(J_t(M_1)-J_t(M_0))| \leq 2\varepsilon \). This is a finite utility perturbation bound, not a training or distribution-shift theorem.
Example 2.4. Equal aggregate score can hide a slice reversal [ftip-00G6]AGENTDRAFTED
Example 2.4. Equal aggregate score can hide a slice reversal [ftip-00G6]AGENTDRAFTED
On equally likely outcomes \(a,b\), take score vectors \((s_{M_1}(a),s_{M_1}(b))=(1,0)\) and \((s_{M_0}(a),s_{M_0}(b))=(0,1)\); a second score map swaps coordinates. Both artifacts have aggregate score \(1/2\) under both maps, but their per-outcome ordering reverses. Aggregate transport alone does not identify localized behavior.
Remark 2.5. Common domains are a transport hypothesis [ftip-00G7]AGENTDRAFTED
Remark 2.5. Common domains are a transport hypothesis [ftip-00G7]AGENTDRAFTED
The bound in Theorem 2.2 requires the same finite outcome support. Changing prompts, parsers, evaluators, or inference budgets is a new law and must be recorded rather than hidden inside a score perturbation.
3. Prompt and research-state transport [ftip-00G8]AGENTDRAFTED
3. Prompt and research-state transport [ftip-00G8]AGENTDRAFTED
Definition 3.1. Research-state observable [ftip-00G9]AGENTDRAFTED
Definition 3.1. Research-state observable [ftip-00G9]AGENTDRAFTED
For a finite record space \(\mathcal R\), an observable is a map \(h:\mathcal R\to \mathcal H\). A summary-only evaluator sees \(h(r)\) and not the underlying record \(r\); the omitted coordinates are unavailable to that evaluator.
Definition 3.2. Compaction kernel [ftip-00GA]AGENTDRAFTED
Definition 3.2. Compaction kernel [ftip-00GA]AGENTDRAFTED
A compaction kernel is a randomized map \(K(dh\mid r)\) from records in \(\mathcal R\) to summaries in \(\mathcal H\). A deterministic summary is the special case \(K(dh\mid r)=\delta _{h(r)}\).
Theorem 3.3. Summary-only indistinguishability [ftip-00GB]AGENTDRAFTED
Theorem 3.3. Summary-only indistinguishability [ftip-00GB]AGENTDRAFTED
Let \(r,r'\in \mathcal R\) induce the same summary law under the kernel in Definition 3.2. Any randomized decision rule that receives only that summary has the same output law in the two worlds.
Proof.
Proof.
The output law is the pushforward of the common summary law through the same decision kernel. Pushforwards of equal finite measures are equal.
Example 3.4. An omitted feasibility constraint [ftip-00GC]AGENTDRAFTED
Example 3.4. An omitted feasibility constraint [ftip-00GC]AGENTDRAFTED
Two research records can share every summary score while one contains a withdrawn lemma and the other contains a feasible proof plan. A summary-only selector therefore chooses identically, although the correct next action differs. This is a finite witness to information loss, not a claim about a particular model.
Remark 3.5. Long-horizon records are empirical transport evidence [ftip-00GD]AGENTDRAFTED
Remark 3.5. Long-horizon records are empirical transport evidence [ftip-00GD]AGENTDRAFTED
The case study in [li2026longhorizon], Sections 4--7, reports file-based memory, human steering, and research-state failures under one system and task. It motivates explicit observables and omitted constraints; it does not supply the finite theorem in Theorem 3.3 or a general model capability law.
4. Evaluation cells and budget accounting [ftip-00GE]AGENTDRAFTED
4. Evaluation cells and budget accounting [ftip-00GE]AGENTDRAFTED
Definition 4.1. Evaluation cell [ftip-00GF]AGENTDRAFTED
Definition 4.1. Evaluation cell [ftip-00GF]AGENTDRAFTED
An evaluation cell is \(C=(M,\mathsf Q,\mathsf I,b,\Lambda ,v)\), where \(M\) is an artifact, \(\mathsf Q\) a task law, \(\mathsf I\) an inference protocol, \(b\) a budget, \(\Lambda \) a seed kernel, and \(v\) an evaluator version. A cell comparison may differ in named coordinates only.
Lemma 4.2. Additive evaluation cost [ftip-00GG]AGENTDRAFTED
Lemma 4.2. Additive evaluation cost [ftip-00GG]AGENTDRAFTED
For a finite evaluation plan with disjoint stages of costs \(c_1,\ldots ,c_k\in \mathbb R_{\geq 0}\), total cost is \(c=\sum _{i=1}^k c_i\). Appending a stage increases cost by its declared amount and cannot preserve a budget \(B\) unless the remaining slack covers it.
Proof.
Proof.
The statement is the associativity and monotonicity of finite sums.
Theorem 4.3. Budget-preserving matched comparison [ftip-00GH]AGENTDRAFTED
Theorem 4.3. Budget-preserving matched comparison [ftip-00GH]AGENTDRAFTED
Let two evaluation plans use the same declared stages except for a substitution whose cost is at most the replaced stage cost. If their common prefix cost is \(c_0\) and both suffix costs fit \(B-c_0\), then both cells in Definition 4.1 are admissible under budget \(B\).
Proof.
Proof.
Apply Lemma 4.2 to the common prefix and each suffix. The assumed inequalities give total cost at most \(B\) in each plan.
Example 4.4. Cost-balanced evaluation cells [ftip-00GI]AGENTDRAFTED
Example 4.4. Cost-balanced evaluation cells [ftip-00GI]AGENTDRAFTED
A paired run can spend \(B=100\) units as 40 units of inference, 20 of evaluator calls, and 40 of audit. Replacing 10 inference units by 10 audit units preserves the budget, but changes the cell and must not be described as the same inference protocol.
Remark 4.5. Execution and evaluator confounds [ftip-00GJ]AGENTDRAFTED
Remark 4.5. Execution and evaluator confounds [ftip-00GJ]AGENTDRAFTED
A score change can arise from the artifact, prompt law, inference budget, parser, evaluator version, or execution failure. A matched-cell report names these coordinates before interpreting a paired gain.
5. Distribution shift and ranking reversals [ftip-00GK]AGENTDRAFTED
5. Distribution shift and ranking reversals [ftip-00GK]AGENTDRAFTED
Example 5.1. Train and evaluation laws can reverse a ranking [ftip-00GL]AGENTDRAFTED
Example 5.1. Train and evaluation laws can reverse a ranking [ftip-00GL]AGENTDRAFTED
Let \(\mathsf Q_{\rm tr}\) put masses \(.9,.1\) on \(a,b\), and let \(\mathsf Q_{\rm ev}\) swap them. Artifacts \(M_1,M_0\) with score vectors \((s_{M_1}(a),s_{M_1}(b))=(1,0)\) and \((s_{M_0}(a),s_{M_0}(b))=(0,1)\) rank oppositely under the two laws. A training score is not an evaluation transport certificate.
Theorem 5.2. Total-variation transport bound [ftip-00GM]AGENTDRAFTED
Theorem 5.2. Total-variation transport bound [ftip-00GM]AGENTDRAFTED
For any score \(s:\mathcal Z\to [0,1]\) and finite laws \(\mathsf Q,\mathsf Q'\), \(|\mathbb E_{\mathsf Q}s-\mathbb E_{\mathsf Q'}s|\leq \lVert \mathsf Q-\mathsf Q'\rVert _{\rm TV}\), where \(\lVert \mathsf Q-\mathsf Q'\rVert _{\rm TV}:=\frac 12\sum _{z\in \mathcal Z} |\mathsf Q(z)-\mathsf Q'(z)|\).
Proof.
Proof.
Expand the finite expectation difference and use \(|s(z)|\leq 1\). The positive and negative parts are bounded by the total variation mass, giving the displayed inequality.
Example 5.3. A sharp two-point transport bound [ftip-00GN]AGENTDRAFTED
Example 5.3. A sharp two-point transport bound [ftip-00GN]AGENTDRAFTED
On \(\mathcal Z=\{a,b\}\), let \(s(a)=1,s(b)=0\), and let \(\mathsf Q(a)=1\), \(\mathsf Q'(a)=1-\eta \), and \(\mathsf Q'(b)=\eta \). The score difference and total variation distance are both \(\eta \).
Remark 5.4. What transport does not identify [ftip-00GO]AGENTDRAFTED
Remark 5.4. What transport does not identify [ftip-00GO]AGENTDRAFTED
The bounds above transport a declared score across a declared law. They do not identify why an artifact changed, establish support expansion, prove generalization, or order different evaluators whose outcome spaces differ.
Example 5.5. A minimal evaluation trace [ftip-00GP]AGENTDRAFTED
Example 5.5. A minimal evaluation trace [ftip-00GP]AGENTDRAFTED
A reproducible record is \((M,\mathsf Q,\mathsf I,b,\Lambda ,v,m, \widehat \Delta _m,c)\), listing the cells, sample count, paired estimate, and cost. Replaying this tuple reproduces the estimator target only when the recorded laws and versions remain available.
Remark 5.6. Execution controls for a matched evaluation [ftip-00GQ]AGENTDRAFTED
Remark 5.6. Execution controls for a matched evaluation [ftip-00GQ]AGENTDRAFTED
A matched evaluation requires execution gates, contamination checks, grader calibration, and an audit of the result before it is committed. These controls make the comparison auditable; they do not by themselves establish capability acquisition.