Remark. RLVR sampling, feedback, and update rules [ftip-0045]
AGENTDRAFTED
Concrete RLVR methods include the PPO-based construction in
[lambert2024tulu3, Section 6], group relative policy optimization (GRPO) in
[shao2024deepseekmath, Section 4.1], R1-Zero in
[guo2025deepseek, Section 2.2], and Decoupled Clip and Dynamic sAmpling
Policy Optimization (DAPO) in
[yu2025dapo, Sections 2--3].
A proposed common description groups methods that acquire policy samples, evaluate a
declared verifiable signal, assign credit, and update the policy. A
specific claim must still name its verifier, group sampling, estimator,
clipping, normalization, reference-policy treatment, and length rule.