Definition. Overlong reward shaping [yu2025dapo, Section 3.4] [ftip-004K]
Definition. Overlong reward shaping [yu2025dapo, Section 3.4] [ftip-004K]
Overlong reward shaping introduces a soft penalty near the maximum response length before the hard truncation boundary. The penalty increases over a declared buffer region rather than assigning the same abrupt terminal penalty to every response that reaches the limit.
The construction changes both the reward value and which length-related behavior receives gradient. The maximum length, buffer width, and penalty schedule are part of the reward design.