Remark. An observed plateau is not a universal probability threshold [ftip-00A7]

[clay2026demystifying, Section 5.2 and Figure 2] report that the Qwen3-1.7B Base movie-quote cell, whose initial empirical match rate is about 0.5 percent, does not optimize under the source's sparse-reward run, whereas its dense-reward run rises toward fifty percent. [clay2026demystifying, Appendix Figure 8] gives final match rates 10.0 percent for sparse reward and 48.8 percent for dense reward. Thus ``failed'' here means failure of the reported sparse run to optimize as intended, not zero final matches.

A plateau is indexed by the model, target, reward implementation, optimizer, sampling process, and budget in its experiment cell. One observed transition therefore does not identify a universal initial-probability threshold for RL learning, support, or capability acquisition.