Remark. What the starting-policy comparison can and cannot identify [ftip-009W]

Within a declared model, task, training grid, and evaluation procedure, the three-cell comparison can establish that the measured post-training outcome differs across the source's constructed starting checkpoints. The movie-quote and AIME entries in [clay2026demystifying, Section 5.1, Tables 1--2] are empirical records of that form.

The comparison does not isolate initial target probability from every other SFT-induced weight change, prove that an observed zero has zero support, or establish acquisition of a broader capability. The reported model sets differ across the paper: [clay2026demystifying, Section 4.1] names OLMo 3, Qwen2-7B, Qwen3-1.7B, and Qwen2.5-7B-Instruct, while Table 1 and Appendix G also report Qwen3-8B and Qwen2-1.5B. Each measurement is specific to the model and configuration in its table row.