Corollary. Sparse-reward discovery within a rollout budget [ftip-00A5]
Corollary. Sparse-reward discovery within a rollout budget [ftip-00A5]
Suppose the events \(\{r_j^{\mathrm {sp}}>0\}\) are independent and have a common probability \(q\). For every positive rollout budget \(B\),
\[ \Pr (J_+\leq B)=1-(1-q)^B. \]This is a specialization proved from the displayed hypotheses. It does not supply independence, stationarity, or the value of \(q\) for a training run, and it does not equate positive proxy reward with task utility.