Remark. GRPO variants and estimator conventions [shao2024deepseekmath, Section 4.1] [ftip-004G]

GRPO names a family of group-relative policy updates rather than one fully determined estimator. Outcome and process supervision assign different reward structures. Implementations may also differ in token versus response normalization, KL terms, clipping intervals, importance weights, and treatment of groups with zero reward variance.

A result about one variant must retain its producing policy and all of those conventions. Sharing the group-relative baseline does not make the resulting gradients identical.