Remark. Verifier error under an independent audit law [ftip-0061]
Remark. Verifier error under an independent audit law [ftip-0061]
Verifier error is measured against an independent success label because a training reward cannot audit itself. The verifier and evidence interfaces are in Definition [ftip-0046]--Definition [ftip-0049]. DeepSeek-Prover-V2 [ren2025deepseekproverv2, Section 2.3, ``Reinforcement Learning''] supplies a proof-checking instance. The false-accept/false-reject decomposition is a binary measurement model.
The rates may vary by task, policy, trajectory length, and adversarial pressure. A low average error rate need not prevent reward hacking on the subpopulation favored by optimization.