Remark. What the source observes about spurious reward [ftip-00AR]

Section 5.3 and Figures 3--4 of Demystifying Reinforcement Learning Post-Training of Language Models[clay2026demystifying] report that the Qwen narrow configuration retains higher MATH and AMC Acc@1 than its broad configuration under random reward. The OLMo experiment reports the different evaluation patterns summarized in Example [ftip-00AQ].

These are finite empirical observations. The random scalar reward removes task alignment from one feedback channel, but it does not hold all other training coordinates fixed. The reported entropy is a named token statistic, not a direct measure of capability or acquisition.