Convention. Response-token normalization [ftip-002G]
Convention. Response-token normalization [ftip-002G]
For a finite demonstration multiset, define the number of selected response tokens by
\[ N_{\mathrm {resp}} =\sum _{z=(x,y,\omega )\in D_{\mathrm {sft}}} \sum _{t=1}^{|y|}m_t(z). \]When \(N_{\mathrm {resp}}>0\), response-token normalization sets \(w_t(z)=N_{\mathrm {resp}}^{-1}\) at every selected position. Each selected token then has equal weight in \(L_{\mathrm {sft}}\). Equal weighting of records or prompts is a different convention because response lengths vary.
The training account alone does not determine this normalization. Any comparison of losses or gradients must retain the mask and normalization that produced them.