Reliability protocol and failure taxonomy [ftip-00GR]
✍️sourceAGENTDRAFTED
Reliability protocol and failure taxonomy [ftip-00GR]
✍️sourceAGENTDRAFTED
A reliability protocol binds an evaluation to its durable state, replay inputs, and audit record. Execution and grading failures can invalidate a measurement even when the reported benchmark score improves.
1. Measurement and contamination controls [ftip-00GS]AGENTDRAFTED
1. Measurement and contamination controls [ftip-00GS]AGENTDRAFTED
Convention 1.1. Evidence required for a reliability claim [ftip-00GT]AGENTDRAFTED
Convention 1.1. Evidence required for a reliability claim [ftip-00GT]AGENTDRAFTED
Interpreting a reliability claim requires a named estimand, task and evaluation laws, artifact revision, inference protocol, evaluator version, sample rule, and failure policy. Missing coordinates are unknown rather than implicitly held fixed.
Definition 1.2. Contamination record [ftip-00GU]AGENTDRAFTED
Definition 1.2. Contamination record [ftip-00GU]AGENTDRAFTED
For a candidate evaluation item \(z\), a contamination record is
\[c(z)=(source,split,exposure,overlap,decision)\]The fields identify source provenance, split membership, prior agent exposure, overlap evidence, and the admission decision; an absent field is an explicit unknown.
Definition 1.3. A contamination predicate [ftip-00GV]AGENTDRAFTED
Definition 1.3. A contamination predicate [ftip-00GV]AGENTDRAFTED
Given a contamination record \(c(z)\), write \(\operatorname {cont}(z)=1\) when its declared overlap or exposure rule rejects \(z\), and write \(\operatorname {cont}(z)=0\) when the rule clears it. The predicate belongs to the protocol, not to a model’s score.
Definition 1.4. An uncontaminated evaluation slice [ftip-00GW]AGENTDRAFTED
Definition 1.4. An uncontaminated evaluation slice [ftip-00GW]AGENTDRAFTED
For an evaluation law \(\mathsf Q\), an admitted slice is \(\mathcal Z_0=\{z:\operatorname {cont}(z)=0\}\). Its reported score is conditioned on the declared slice law; removing contaminated items does not preserve the original estimand unless the law is unchanged by construction.
Definition 1.5. A paired comparison design [ftip-00GX]AGENTDRAFTED
Definition 1.5. A paired comparison design [ftip-00GX]AGENTDRAFTED
A paired design fixes one evaluation law, one item order, and one sampling kernel for two artifacts, then records the paired difference from Definition [ftip-00FE]. Any item-level exclusion is applied before seeing the paired outcomes or is declared as an adaptive rule.
Lemma 1.6. Pairing preserves the declared marginal target [ftip-00GY]AGENTDRAFTED
Lemma 1.6. Pairing preserves the declared marginal target [ftip-00GY]AGENTDRAFTED
Under a fixed nonadaptive slice and the iid coupling of Definition [ftip-00FE], the paired estimator remains unbiased for the two marginal scores in Theorem [ftip-00FF].
Proof.
Proof.
Condition on the fixed slice. The coupling has the stated marginals, so the proof of Theorem [ftip-00FF] applies without changing either expectation.
Example 1.7. A contamination decision can change the estimand [ftip-00GZ]AGENTDRAFTED
Example 1.7. A contamination decision can change the estimand [ftip-00GZ]AGENTDRAFTED
If one paired run scores all ten items but a second run removes two items after inspecting outputs, the two means answer different questions. A report must expose the removal rule and the resulting slice law.
Remark 1.8. Scope of measurement methodology [ftip-00H0]AGENTDRAFTED
Remark 1.8. Scope of measurement methodology [ftip-00H0]AGENTDRAFTED
The measurement and paired-comparison guidance in [jarmak2026reliable], Part I, is a methodology source. It motivates contamination ledgers and matched comparisons; it supplies no universal contamination detector or statistical guarantee.
2. Grader calibration [ftip-00H1]AGENTDRAFTED
2. Grader calibration [ftip-00H1]AGENTDRAFTED
Definition 2.1. A typed grading rubric [ftip-00H2]AGENTDRAFTED
Definition 2.1. A typed grading rubric [ftip-00H2]AGENTDRAFTED
A grading rubric is \(R=(\mathcal Z,\mathcal L,g,v)\): outcome space, finite label set, scoring map \(g:\mathcal Z\to \mathcal L\), and evaluator version \(v\). The rubric declares which labels count as success and which observations are abstentions.
Definition 2.2. A grader response law [ftip-00H3]AGENTDRAFTED
Definition 2.2. A grader response law [ftip-00H3]AGENTDRAFTED
For rubric \(R\), a grader response law is a kernel \(K(\ell \mid z)\) on \(\mathcal L\) given outcome \(z\). A deterministic grader is the point-mass case; stochastic or model-based graders must retain their version and sampling coordinates.
Definition 2.3. A calibration set [ftip-00H4]AGENTDRAFTED
Definition 2.3. A calibration set [ftip-00H4]AGENTDRAFTED
A calibration set is a finite set \(\mathcal C\subseteq \mathcal Z\) with an adjudicated label \(y^star(z)\) for each \(z\in \mathcal C\). The adjudication source, disagreement policy, and release version are part of the set’s provenance.
Definition 2.4. Calibration agreement [ftip-00H5]AGENTDRAFTED
Definition 2.4. Calibration agreement [ftip-00H5]AGENTDRAFTED
For grader kernel \(K\) and calibration set \(\mathcal C\), the agreement rate is \(A(K,\mathcal C)=|\mathcal C|^{-1}\sum _{z\in \mathcal C} K(y^star(z)\mid z)\). It is a calibration-set statistic, not a population accuracy claim.
Lemma 2.5. Agreement threshold admission [ftip-00H6]AGENTDRAFTED
Lemma 2.5. Agreement threshold admission [ftip-00H6]AGENTDRAFTED
If \(|\mathcal C|=n\geq 1\) and \(A(K,\mathcal C)\geq 1-\eta \), then the empirical disagreement mass \(1-A(K,\mathcal C)\) is at most \(\eta \).
Proof.
Proof.
Rearrange the defining finite average in Definition 2.4. No generalization beyond \(\mathcal C\) is implied.
Example 2.6. High aggregate agreement can hide a subgroup failure [ftip-00H7]AGENTDRAFTED
Example 2.6. High aggregate agreement can hide a subgroup failure [ftip-00H7]AGENTDRAFTED
A grader that agrees on 99 of 100 easy cases but disagrees on the one safety case has overall agreement \(99/101\approx 0.9802\). Its agreement within the safety subgroup is zero. Any positive safety-subgroup agreement threshold therefore rejects this grader despite its high overall agreement. A calibration report stratifies by failure-relevant labels.
Definition 2.7. A grader release gate [ftip-00H8]AGENTDRAFTED
Definition 2.7. A grader release gate [ftip-00H8]AGENTDRAFTED
A grader release gate accepts \(K\) only when its calibration version, agreement thresholds, abstention handling, and stratified failure checks are recorded. A failed gate blocks interpretation of downstream score changes.
Remark 2.8. Scope of grading methodology [ftip-00H9]AGENTDRAFTED
Remark 2.8. Scope of grading methodology [ftip-00H9]AGENTDRAFTED
Part II of [jarmak2026reliable] organizes grading and reviewer calibration practices. The finite agreement quantities do not establish a universal threshold or blinded-human reliability theorem.
3. Execution gates and durable state [ftip-00HA]AGENTDRAFTED
3. Execution gates and durable state [ftip-00HA]AGENTDRAFTED
Definition 3.1. An execution gate [ftip-00HB]AGENTDRAFTED
Definition 3.1. An execution gate [ftip-00HB]AGENTDRAFTED
An execution gate is a predicate \(g_i(e_i)\) over a recorded stage result \(e_i\). It names its owner, required inputs, pass/fail/unknown outputs, and the artifact versions to which the result applies.
Definition 3.2. An ordered gate chain [ftip-00HC]AGENTDRAFTED
Definition 3.2. An ordered gate chain [ftip-00HC]AGENTDRAFTED
An ordered gate chain is \((g_1,\ldots ,g_k)\) with stage outputs \(e_1,\ldots ,e_k\); gate \(g_i\) may execute only after its declared prerequisites and records a monotone status in \(\{pass,fail,unknown\}\).
Theorem 3.3. A failed gate blocks a release conjunction [ftip-00HD]AGENTDRAFTED
Theorem 3.3. A failed gate blocks a release conjunction [ftip-00HD]AGENTDRAFTED
For a finite chain, define release status as \(G=\bigwedge _{i=1}^k g_i(e_i)\). If any gate is false, then \(G\) is false.
Proof.
Proof.
This is the defining conjunction of finitely many Boolean gate predicates. It says nothing about whether a later gate would have passed.
Definition 3.4. A durable execution record [ftip-00HE]AGENTDRAFTED
Definition 3.4. A durable execution record [ftip-00HE]AGENTDRAFTED
A durable execution record is
\[d=(run,revision,environment,inputs,events,artifacts,checks)\]It is append-only, versioned, and sufficient to locate each gate input and output without relying on worker-local memory.
Definition 3.5. A replay contract [ftip-00HF]AGENTDRAFTED
Definition 3.5. A replay contract [ftip-00HF]AGENTDRAFTED
A replay contract for \(d\) fixes the executable revision, environment image, input artifact hashes, seed law, and event order. A replay is faithful only when those coordinates are available and the resulting events satisfy the recorded schema.
Theorem 3.6. Deterministic replay reproduces a recorded trace [ftip-00HG]AGENTDRAFTED
Theorem 3.6. Deterministic replay reproduces a recorded trace [ftip-00HG]AGENTDRAFTED
If the execution map is deterministic in the coordinates fixed by Definition 3.5, replaying the same inputs, revision, environment, and seed produces the same event trace.
Proof.
Proof.
Induct over the finite event sequence. Equal initial coordinates give equal first events; determinism and equal prefixes give the next event.
Example 3.7. Replayability is not correctness [ftip-00HH]AGENTDRAFTED
Example 3.7. Replayability is not correctness [ftip-00HH]AGENTDRAFTED
A reproducible run can deterministically reproduce a wrong patch or a misconfigured evaluator. Replay establishes trace identity under its contract; it does not establish that the trace met the intended utility or safety goal.
Definition 3.8. A state-integrity digest [ftip-00HI]AGENTDRAFTED
Definition 3.8. A state-integrity digest [ftip-00HI]AGENTDRAFTED
For durable record \(d\), let \(H(d)\) be a cryptographic digest of its canonical serialized fields. A replay request fails closed when the supplied record digest or any referenced input hash differs from the declared value.
To bind a replay result to an audit decision, the replay trace must be included in the audited record and covered by its digest. A matching digest for other fields does not bind that trace to the decision.
Remark 3.9. Scope of execution methodology [ftip-00HJ]AGENTDRAFTED
Remark 3.9. Scope of execution methodology [ftip-00HJ]AGENTDRAFTED
Part III of [jarmak2026reliable] motivates containment, durable execution, and recovery records. Deterministic replay and matching digests establish execution consistency; they do not ensure correctness.
4. Audit before commit and failure taxonomy [ftip-00HK]AGENTDRAFTED
4. Audit before commit and failure taxonomy [ftip-00HK]AGENTDRAFTED
Definition 4.1. An audit-before-commit record [ftip-00HL]AGENTDRAFTED
Definition 4.1. An audit-before-commit record [ftip-00HL]AGENTDRAFTED
An audit-before-commit record is \(a=(d,scope,checks,reviewer,decision)\). It binds the durable execution record \(d\) to the audited scope, check list, review identity, and a decision in \(\{commit,hold,reject\}\).
Theorem 4.2. Audit-before-commit safety gate [ftip-00HM]AGENTDRAFTED
Theorem 4.2. Audit-before-commit safety gate [ftip-00HM]AGENTDRAFTED
A commit decision is admissible only if every required check in an Definition 4.1 record is pass, the audited digest equals the proposed record digest, and the decision is \(commit\).
Proof.
Proof.
Admissibility is defined as the conjunction of these three recorded conditions; failure of any conjunct yields hold or reject.
Definition 4.3. An independent audit scope [ftip-00HN]AGENTDRAFTED
Definition 4.3. An independent audit scope [ftip-00HN]AGENTDRAFTED
An audit is independent for scope \(S\) when its inputs are frozen before review, its reviewer or checker is distinct from the proposing action, and its decision cannot rewrite \(S\) without a new record. Independence is a protocol condition, not a claim of infallibility.
Example 4.4. A latent failure caught before commit [ftip-00HO]AGENTDRAFTED
Example 4.4. A latent failure caught before commit [ftip-00HO]AGENTDRAFTED
A run may pass unit tests while an independent audit finds that the evaluator version in \(d\) differs from the declared version. The audit holds the commit and records the mismatch; replay alone would not have caught it.
Definition 4.5. A reliability failure taxonomy [ftip-00HP]AGENTDRAFTED
Definition 4.5. A reliability failure taxonomy [ftip-00HP]AGENTDRAFTED
Classify a failed run by its earliest violated contract: measurement failure (M), grading failure (G), execution/state failure (E), audit failure (A), or transfer failure (T). A single run may receive secondary labels, but the primary label is the earliest failed gate in recorded order.
Lemma 4.6. Earliest-failure labels are unique [ftip-00HQ]AGENTDRAFTED
Lemma 4.6. Earliest-failure labels are unique [ftip-00HQ]AGENTDRAFTED
For a finite ordered gate chain with at least one failed gate, the earliest failed index is unique and therefore determines one primary taxonomy label.
Proof.
Proof.
Every nonempty finite subset of ordered indices has a unique least element.
Definition 4.7. A remediation record [ftip-00HR]AGENTDRAFTED
Definition 4.7. A remediation record [ftip-00HR]AGENTDRAFTED
A remediation record maps a primary label from Definition 4.5 to an owner, corrective action, re-test scope, and closure condition. Closure requires a new durable record and cannot mutate the failed record in place.
Remark 4.8. Scope of audit methodology [ftip-00HS]AGENTDRAFTED
Remark 4.8. Scope of audit methodology [ftip-00HS]AGENTDRAFTED
Parts III and V of [jarmak2026reliable] motivate durable execution, review, and accountability practices. The audit taxonomy and conjunction gates are proposed operational specifications; they do not establish that an audit is complete or that a committed artifact is correct.
5. Audit budgets and the scope of run-level evidence [ftip-00HT]AGENTDRAFTED
5. Audit budgets and the scope of run-level evidence [ftip-00HT]AGENTDRAFTED
Remark 5.1. A protocol is not a causal estimator [ftip-00HU]AGENTDRAFTED
Remark 5.1. A protocol is not a causal estimator [ftip-00HU]AGENTDRAFTED
The gates above make evidence auditable. They do not identify the causal effect of an intervention when task law, artifact, grader, budget, or harness state changes together. Those coordinates must be controlled in the matched evaluation of Definition [ftip-00GF].
Example 5.2. Counterexample: a perfect replay can repeat a contaminated result [ftip-00HV]AGENTDRAFTED
Example 5.2. Counterexample: a perfect replay can repeat a contaminated result [ftip-00HV]AGENTDRAFTED
Let a deterministic run train and evaluate on one leaked item. Its digest, event trace, and replay are all identical across reruns, yet the measurement is contaminated under Definition 1.3. Replayability and contamination control are independent obligations.
Example 5.3. Counterexample: a calibrated grader can still face a changed domain [ftip-00HW]AGENTDRAFTED
Example 5.3. Counterexample: a calibrated grader can still face a changed domain [ftip-00HW]AGENTDRAFTED
A grader can meet its threshold on \(\mathcal C\) in Lemma 2.5 while every deployment item lies outside that calibration domain and receives an unvalidated label. Agreement on \(\mathcal C\) alone does not transport to a new law.
Theorem 5.4. Budgeted audit admission [ftip-00HX]AGENTDRAFTED
Theorem 5.4. Budgeted audit admission [ftip-00HX]AGENTDRAFTED
If an audit plan has finite nonnegative stage costs and its declared sum is at most budget \(B\), then the plan is budget-admissible; adding any positive stage beyond the remaining slack makes it inadmissible.
Proof.
Proof.
Apply additive cost monotonicity from Lemma [ftip-00GG].
Example 5.5. A cost-balanced audit plan [ftip-00HY]AGENTDRAFTED
Example 5.5. A cost-balanced audit plan [ftip-00HY]AGENTDRAFTED
With budget \(B=100\), a plan may allocate 50 units to execution, 30 to grading calibration, and 20 to audit. A proposed 10-unit extra audit must be funded by reducing another stage or the admission gate fails.
Remark 5.6. What this protocol can establish [ftip-00HZ]AGENTDRAFTED
Remark 5.6. What this protocol can establish [ftip-00HZ]AGENTDRAFTED
A passing protocol establishes that declared evidence contracts, hashes, checks, and budgets were satisfied for the recorded run. It does not establish capability acquisition, broad generalization, or absence of unobserved faults.