Remark. Normalization of masked pretraining loss [ftip-006L]

The autoregressive token loss follows [phuong2022formal, Section 7]. Kaplan et al. report language-model loss with parameter and compute variables in [kaplan2020scaling, sec. 2.1]. The population ratio-of-expectations in Definition [ftip-001C] and finite per-token ratio in Definition [ftip-001E] are proposed normalization conventions with explicit prediction masks. Neither cited source defines those exact ratios.

A theorem or experiment using another record weighting, length weighting, or expectation order must state that change rather than reuse these symbols.