Evaluation laws and couplings [ftip-00FA]
✍️sourceAGENTDRAFTED
Evaluation laws and couplings [ftip-00FA]
✍️sourceAGENTDRAFTED
Convention 1. A named evaluation experiment [ftip-00FB]AGENTDRAFTED
Convention 1. A named evaluation experiment [ftip-00FB]AGENTDRAFTED
A named evaluation experiment records the artifact, task law, inference protocol, evaluator, seed kernel, budget, and sampling unit. Two reported scores are comparable only after these coordinates and their intended target have been stated.
Definition 2. Evaluation law [ftip-00FC]AGENTDRAFTED
Definition 2. Evaluation law [ftip-00FC]AGENTDRAFTED
For a fixed evaluation interface from Convention [ftip-005D], an evaluation law is a probability measure \(\mathsf Q\) on the finite outcome space \(\mathcal Z=\mathcal X_{\mathsf T_{\rm ev}}\times \mathcal O_{\mathsf T_{\rm ev}}\) with random pair \((X,O)\sim \mathsf Q\). Its score for artifact \(M\) is a bounded measurable map \(s_M:\mathcal Z\to [0,1]\). Write \(J_{\mathsf Q}(M)=\mathbb E_{\mathsf Q}[s_M(X,O)]\).
Definition 3. Matched evaluation pair [ftip-00FD]AGENTDRAFTED
Definition 3. Matched evaluation pair [ftip-00FD]AGENTDRAFTED
A matched evaluation pair is \((\mathsf Q,\mathsf I,b,\Lambda )\) and \((\mathsf Q',\mathsf I',b',\Lambda ')\) together with a coupling \((Z,Z')\) whose marginals are the two laws. The pair is matched on the declared coordinates when the task, utility, evaluator version, and sampling unit are shared; any changed coordinate is named.
Definition 4. Paired finite score estimator [ftip-00FE]AGENTDRAFTED
Definition 4. Paired finite score estimator [ftip-00FE]AGENTDRAFTED
For iid coupled draws \((Z_j,Z'_j)_{j=1}^m\), define \(\widehat \Delta _m=m^{-1}\sum _{j=1}^m(s_M(Z_j)-s_{M'}(Z'_j))\). The pairing is part of the estimator specification; it does not silently identify the two evaluation laws.
Theorem 5. Unbiasedness of the paired score estimator [ftip-00FF]AGENTDRAFTED
Theorem 5. Unbiasedness of the paired score estimator [ftip-00FF]AGENTDRAFTED
Under the iid coupling in Definition 4, \(\mathbb E[\widehat \Delta _m]=J_{\mathsf Q}(M)-J_{\mathsf Q'}(M')\).
Proof.
Proof.
Linearity of expectation gives the displayed identity.
\[ \mathbb E[\widehat \Delta _m] =m^{-1}\sum _j\left (\mathbb E[s_M(Z_j)]- \mathbb E[s_{M'}(Z'_j)]\right ). \]The coupling fixes the joint sampling but does not change either marginal, so the two terms are the stated expectations.
Example 6. Common seeds can reduce paired variance [ftip-00G0]AGENTDRAFTED
Example 6. Common seeds can reduce paired variance [ftip-00G0]AGENTDRAFTED
If \(s_M(Z)=s_{M'}(Z')\) almost surely under a common-seed coupling, then \(\widehat \Delta _m=0\) for every sample even when independent sampling would have nonzero variance. This is a variance statement, not evidence that the two artifacts have equal behavior on a different law.
Remark 7. A coupling does not merge estimands [ftip-00G1]AGENTDRAFTED
Remark 7. A coupling does not merge estimands [ftip-00G1]AGENTDRAFTED
The result in Theorem 5 follows from its displayed hypotheses. A shared seed can make a paired comparison precise while the changed task law, evaluator, or inference budget still changes the estimand. No causal effect follows without a declared intervention and matched coordinates.