Prompt breadth and effect locality [ftip-00AK]
✍️sourceAGENTDRAFTED
Prompt breadth and effect locality [ftip-00AK]
✍️sourceAGENTDRAFTED
Training on one prompt law can change behavior unevenly across evaluation slices. Localized and distributed degradation are distinguished relative to that training law and the specified evaluation slices.
Definition 1. Training prompt law and evaluated prompt slice [ftip-00AL]AGENTDRAFTED
Definition 1. Training prompt law and evaluated prompt slice [ftip-00AL]AGENTDRAFTED
Let \(\mathcal X\) be a task-instance set. Write \(\mu _{\rm tr}=\mathcal D_{\rm tr}\) for the training-prompt coordinate of an experiment cell in Definition [ftip-009T]. A training prompt law is this probability law on \(\mathcal X\), used to draw instances during a post-training protocol. Let \(\mu _{\rm ev}\) be an independently declared evaluation law on the same set.
An evaluated prompt slice is a measurable set \(C\subseteq \mathcal X\) with \(\mu _{\rm ev}(C)>0\). Its conditional evaluation law is
\[ \mu _{\rm ev}(A\mid C) =\frac {\mu _{\rm ev}(A\cap C)}{\mu _{\rm ev}(C)}. \]The training and evaluation laws need not agree. A slice records where an effect is measured; it does not assert that prompts inside the slice are equally difficult or represented equally in pretraining. We specialize the task-law convention of Definition [ftip-001Q].
Definition 2. Source-specific narrow and broad configurations [ftip-00AM]AGENTDRAFTED
Definition 2. Source-specific narrow and broad configurations [ftip-00AM]AGENTDRAFTED
A prompt configuration is the record
\[ \mathsf c=(M_0,\mu _{\rm tr},m_{\rm tr},r,\mathsf A,b), \]containing the starting artifact, training prompt law, finite prompt-pool size, reward, update algorithm, and training budget. We call two particular records \(\mathsf c_{\rm nar}\) and \(\mathsf c_{\rm brd}\) only when a cited experiment names them narrow and broad.
Section 4.3 and Appendix F of Demystifying Reinforcement Learning Post-Training of Language Models[clay2026demystifying] instantiate these labels in two model-family experiments. The Qwen comparison uses 100 DeepScaleR prompts versus 10,000 WildChat prompts. The OLMo comparison uses 100 math-only prompts versus 10,000 prompts split evenly among mathematics, instruction following, and code. These labels identify the two reported configurations; they do not define a general measure or ordering of prompt breadth.
Remark 3. Prompt breadth is not a total order [ftip-00AN]AGENTDRAFTED
Remark 3. Prompt breadth is not a total order [ftip-00AN]AGENTDRAFTED
The records in Definition 2 change several coordinates together. The prompt-pool size, domain mixture, and repetition frequency differ. Across the paper's Qwen and OLMo studies, the starting artifact and evaluated tasks also differ.
We therefore use ``narrow'' and ``broad'' as names for the source cells. They do not define a total order on prompt laws, and their comparison does not identify a causal effect of breadth alone. A breadth theorem would need a declared statistic and matched configurations that differ only in that statistic.
Definition 4. Slice-conditioned success change [ftip-00AO]AGENTDRAFTED
Definition 4. Slice-conditioned success change [ftip-00AO]AGENTDRAFTED
Fix an evaluated slice \(C\) from Definition 1, and let \(P_0\) and \(P_1\) denote the pre- and post-training protocols whose successful-support probabilities are defined in Definition [ftip-005U]. The slice-conditioned success change is
\[ \Delta _{\rm suc}(C;P_1,P_0) =\mathbb E_{X\sim \mu _{\rm ev}(\cdot \mid C)} \left [p_{P_1}(X)-p_{P_0}(X)\right ]. \]The evaluation interface, inference budget, and success predicate are held fixed across the two terms. The quantity is an average change on one declared slice. It neither locates the change inside the model nor identifies which training coordinate caused it.
Definition 5. Localized and distributed degradation [ftip-00AP]AGENTDRAFTED
Definition 5. Localized and distributed degradation [ftip-00AP]AGENTDRAFTED
Fix pairwise disjoint evaluated slices \(C_1,\ldots ,C_J\) and a declared degradation threshold \(\kappa >0\). Define the set of materially degraded slices
\[ \mathcal H_\kappa (P_1,P_0) =\left \{j:\Delta _{\rm suc}(C_j;P_1,P_0)\leq -\kappa \right \}. \]Degradation is localized relative to this slice family and threshold when \(|\mathcal H_\kappa |=1\). It is distributed when \(|\mathcal H_\kappa |\geq 2\).
These terms depend on the chosen partition, threshold, and evaluation budget. They do not turn a finite benchmark vector into a statement about all capabilities.
Example 6. Narrow and broad random-reward paths [ftip-00AQ]AGENTDRAFTED
Example 6. Narrow and broad random-reward paths [ftip-00AQ]AGENTDRAFTED
Section 5.3 and Appendix F of Demystifying Reinforcement Learning Post-Training of Language Models[clay2026demystifying] supply two OLMo paths; Figure 4 reports their outcomes. Both replace the usual verifiable reward with a random scalar drawn from \(\operatorname {Unif}[0,1]\), but their prompt configurations differ.
In the narrow SFT-start path, GSM8K accuracy falls from 86 to about 32 by step 400, while the reported MMLU and IFEval changes are much smaller. In the broad path, the source reports an entropy spike near step 400 together with collapse on GSM8K, MMLU, and IFEval. The diagram records that interpretation; it does not isolate prompt breadth from pool size, mixture, or repetition.
Remark 7. What the source observes about spurious reward [ftip-00AR]AGENTDRAFTED
Remark 7. What the source observes about spurious reward [ftip-00AR]AGENTDRAFTED
Section 5.3 and Figures 3--4 of Demystifying Reinforcement Learning Post-Training of Language Models[clay2026demystifying] report that the Qwen narrow configuration retains higher MATH and AMC Acc@1 than its broad configuration under random reward. The OLMo experiment reports the different evaluation patterns summarized in Example 6.
These are finite empirical observations. The random scalar reward removes task alignment from one feedback channel, but it does not hold all other training coordinates fixed. The reported entropy is a named token statistic, not a direct measure of capability or acquisition.
Example 8. Counterexample: equal marginal reward, unequal prompt-local damage [ftip-00AS]AGENTDRAFTED
Example 8. Counterexample: equal marginal reward, unequal prompt-local damage [ftip-00AS]AGENTDRAFTED
Let two evaluation slices \(C_1,C_2\) have equal mass. A feedback channel returns an independent Bernoulli reward with mean \(1/2\) under either of two post-training protocols. Suppose their fixed-interface success changes are
\[ \left (\Delta _{\rm suc}(C_1),\Delta _{\rm suc}(C_2)\right ) =(-1,0) \quad \hbox {or}\quad \left (-\tfrac 12,-\tfrac 12\right ). \]The reward law and its mean are identical, but the slice-level evaluation vectors differ. Hence marginal training reward does not determine whether damage is localized or distributed. This is a finite counterexample, not a claim about the mechanism in the source experiment.
Remark 9. Coverage, feedback resolution, and prompt-conditioned effects [ftip-00AT]AGENTDRAFTED
Remark 9. Coverage, feedback resolution, and prompt-conditioned effects [ftip-00AT]AGENTDRAFTED
The source cells motivate tests of starting-policy coverage, feedback resolution, and prompt-conditioned damage. They do not establish that dense reward creates mathematical support, that zero sampled hits imply zero probability, or that random reward has one model-independent effect.
Separating the effects of coverage, feedback resolution, and prompt breadth requires stronger controls. One comparison fixes prompt laws and varies only the feedback sigma-algebra. Another fixes feedback and evaluation while varying one starting-policy coordinate. Confidence regions for the complete vector of slice-conditioned changes would quantify effects that an aggregate score can hide.
The reported observations remain conditional on their experimental configurations. They do not establish general results about capability acquisition, elicitation, or the optimal allocation of post-training compute.