Reinforcement learning with verifiable rewards [ftip-0042]

Reinforcement learning with verifiable rewards replaces or supplements a learned judge with a declared check of the generated result. Contemporary methods differ in their verifier, sampling law, advantage estimator, clipping, normalization, and treatment of length. This section defines those choices separately before they are assembled into a training protocol.