Definition. Source-specific narrow and broad configurations [ftip-00AM]

A prompt configuration is the record

\[ \mathsf c=(M_0,\mu _{\rm tr},m_{\rm tr},r,\mathsf A,b), \]

containing the starting artifact, training prompt law, finite prompt-pool size, reward, update algorithm, and training budget. We call two particular records \(\mathsf c_{\rm nar}\) and \(\mathsf c_{\rm brd}\) only when a cited experiment names them narrow and broad.

Section 4.3 and Appendix F of Demystifying Reinforcement Learning Post-Training of Language Models[clay2026demystifying] instantiate these labels in two model-family experiments. The Qwen comparison uses 100 DeepScaleR prompts versus 10,000 WildChat prompts. The OLMo comparison uses 100 math-only prompts versus 10,000 prompts split evenly among mathematics, instruction following, and code. These labels identify the two reported configurations; they do not define a general measure or ordering of prompt breadth.