Remark. From a data pipeline to a corpus law [ftip-0018]

The source mixture and cleaning pipeline determine different parts of the data law. The mixture in Definition [ftip-0016] describes how raw documents are proposed. The construction in Definition [ftip-0017] determines which proposals survive and how they are transformed. Together with all random seeds and thresholds, they induce a cleaned document law \(Q_{\mathrm {clean}}\).

Kaplan et al. describe the concrete WebText2 dataset in [kaplan2020scaling, sec. 2.3]. Hoffmann et al. report the MassiveText sources and mixture in [hoffmann2022training, Appendix A, Table A1]. A finite released corpus is one realization of such choices, and an undocumented change in them is not identified by a loss--compute curve.