Remark. Decoding and test-time compute belong to evaluation [ftip-001S]
Remark. Decoding and test-time compute belong to evaluation [ftip-001S]
One next-token law supports many evaluation procedures. Greedy decoding, temperature sampling, majority voting, verifier-guided selection, and multi-round tool use can produce different outcome laws from the same checkpoint. The sampling and selection procedures in [shao2024deepseekmath, secs. 3--4] are concrete examples.
Accordingly, a comparison that changes \(b_{\mathrm {eval}}\) or \(\mathsf I\) does not estimate a training-only effect. A potential frontier therefore depends on evaluation-time compute as well as the trained model.