Controlled coverage, feedback resolution, and prompt breadth [ftip-009M]

This section studies finite measurements and controlled comparisons that separate starting-policy coverage from the information supplied by a reward. It also records how a training prompt law limits the scope of an observed post-training effect.

The empirical route is a controlled language-model post-training study. Its reported outcomes remain experiment-specific observations. Theorems in this section follow from the stated finite-probability assumptions; the empirical results alone do not imply those conclusions.