Remark. The source experiment as an intervention matrix [ftip-009O]

[clay2026demystifying, Sections 4.1--4.3 and Figure 1] organize the reported experiments along three declared axes: a starting model distribution, a training-prompt distribution, and a reward function. Each reported measurement therefore belongs to a cell of a matrix rather than to an unqualified ``RL post-training'' condition. The source holds its broad training recipe fixed while varying selected axes; it does not study variation among RL algorithms.

Each experimental comparison is conditional on its model, task, prompt law, executable reward rule, training budget, and evaluation procedure. The reward rule is especially consequential here because the main text and Appendix C do not state identical sparse-reward semantics. Section 4.2 describes target containment with a length penalty. Appendix C's prose says an exact target receives unit reward, but its displayed equation requires the response to be strictly longer than the target for any positive reward. These two descriptions do not identify a single sparse-reward implementation.