Remark. These interventions change sampling and reward geometry [yu2025dapo, Sections 3.1--3.4] [ftip-004L]

Clip-higher changes the saturated region of the policy surrogate, dynamic sampling conditions which prompt groups enter a batch, token-level aggregation changes response-length weighting, and overlong shaping changes the reward near a truncation boundary. These are four different interventions.

An observed training improvement cannot be assigned to a generic ``algorithm'' without an ablation that holds the other sampling, estimator, reward, and compute choices fixed. The interventions also change which rollouts and gradients are observed, so their effects need not add linearly.