Definition. Outcome-supervision GRPO [shao2024deepseekmath, Section 4.1.2] [ftip-004E]
Definition. Outcome-supervision GRPO [shao2024deepseekmath, Section 4.1.2] [ftip-004E]
Outcome-supervision GRPO samples a group of responses for one prompt, assigns each response an outcome reward, standardizes those rewards within the group, and uses the resulting group-relative advantage across the response's token log-likelihood terms.
The construction removes a separately learned critic but retains a behavior policy, likelihood ratios, a reference-policy term in the source formulation, and explicit normalization choices.