Definition. Token-level policy-gradient loss [yu2025dapo, Section 3.3] [ftip-004J]
Definition. Token-level policy-gradient loss [yu2025dapo, Section 3.3] [ftip-004J]
The token-level loss aggregates active token terms across the complete minibatch and divides by the number of active tokens. A longer response therefore contributes in proportion to its active token count rather than receiving the same total weight as every shorter response.
This normalization changes the empirical gradient even when the sampled responses, rewards, advantages, and likelihood ratios are unchanged. It must be distinguished from averaging a separately normalized loss over responses.