Definition. Empirical pretraining objective [ftip-001E]

For a finite training sample of packed records \(\mathcal D_N=(\zeta ^{(i)})_{i=1}^{N}\) with at least one unmasked target token, the empirical pretraining objective is \[ L_N(\theta ) =\frac { \sum _{i=1}^{N}\sum _{t=2}^{T} \ell _t(\theta ;\zeta ^{(i)}) }{ \sum _{i=1}^{N}\sum _{t=2}^{T}m_t(\zeta ^{(i)}) }. \] This is a per-predicted-token average. Reusing or resampling examples changes the stochastic optimization path even when the displayed finite-sample function is unchanged.