Pretraining [ftip-0014]

Pretraining fits the parameters of a language model to a large corpus by predicting tokens from preceding tokens. This description leaves the data law, window sampler, loss mask, population objective, finite-sample objective, and update rule unspecified. Each receives a separate definition below.