Grader calibration [ftip-00H1]
✍️sourceAGENTDRAFTED
Grader calibration [ftip-00H1]
✍️sourceAGENTDRAFTED
Definition 1. A typed grading rubric [ftip-00H2]AGENTDRAFTED
Definition 1. A typed grading rubric [ftip-00H2]AGENTDRAFTED
A grading rubric is \(R=(\mathcal Z,\mathcal L,g,v)\): outcome space, finite label set, scoring map \(g:\mathcal Z\to \mathcal L\), and evaluator version \(v\). The rubric declares which labels count as success and which observations are abstentions.
Definition 2. A grader response law [ftip-00H3]AGENTDRAFTED
Definition 2. A grader response law [ftip-00H3]AGENTDRAFTED
For rubric \(R\), a grader response law is a kernel \(K(\ell \mid z)\) on \(\mathcal L\) given outcome \(z\). A deterministic grader is the point-mass case; stochastic or model-based graders must retain their version and sampling coordinates.
Definition 3. A calibration set [ftip-00H4]AGENTDRAFTED
Definition 3. A calibration set [ftip-00H4]AGENTDRAFTED
A calibration set is a finite set \(\mathcal C\subseteq \mathcal Z\) with an adjudicated label \(y^star(z)\) for each \(z\in \mathcal C\). The adjudication source, disagreement policy, and release version are part of the set’s provenance.
Definition 4. Calibration agreement [ftip-00H5]AGENTDRAFTED
Definition 4. Calibration agreement [ftip-00H5]AGENTDRAFTED
For grader kernel \(K\) and calibration set \(\mathcal C\), the agreement rate is \(A(K,\mathcal C)=|\mathcal C|^{-1}\sum _{z\in \mathcal C} K(y^star(z)\mid z)\). It is a calibration-set statistic, not a population accuracy claim.
Lemma 5. Agreement threshold admission [ftip-00H6]AGENTDRAFTED
Lemma 5. Agreement threshold admission [ftip-00H6]AGENTDRAFTED
If \(|\mathcal C|=n\geq 1\) and \(A(K,\mathcal C)\geq 1-\eta \), then the empirical disagreement mass \(1-A(K,\mathcal C)\) is at most \(\eta \).
Proof.
Proof.
Rearrange the defining finite average in Definition 4. No generalization beyond \(\mathcal C\) is implied.
Example 6. High aggregate agreement can hide a subgroup failure [ftip-00H7]AGENTDRAFTED
Example 6. High aggregate agreement can hide a subgroup failure [ftip-00H7]AGENTDRAFTED
A grader that agrees on 99 of 100 easy cases but disagrees on the one safety case has overall agreement \(99/101\approx 0.9802\). Its agreement within the safety subgroup is zero. Any positive safety-subgroup agreement threshold therefore rejects this grader despite its high overall agreement. A calibration report stratifies by failure-relevant labels.
Definition 7. A grader release gate [ftip-00H8]AGENTDRAFTED
Definition 7. A grader release gate [ftip-00H8]AGENTDRAFTED
A grader release gate accepts \(K\) only when its calibration version, agreement thresholds, abstention handling, and stratified failure checks are recorded. A failed gate blocks interpretation of downstream score changes.
Remark 8. Scope of grading methodology [ftip-00H9]AGENTDRAFTED
Remark 8. Scope of grading methodology [ftip-00H9]AGENTDRAFTED
Part II of [jarmak2026reliable] organizes grading and reviewer calibration practices. The finite agreement quantities do not establish a universal threshold or blinded-human reliability theorem.