Example. How generation share changes the predicted gain [ftip-00LF]

Keep the illustrative generation factor \(s_{\rm gen}=2.2\) from Example [ftip-00L3] and vary the baseline generation share \(f\). Its cost model gives the following total attempt-throughput factors:

Generation share Predicted throughput factor
20 percent 1.122
50 percent 1.375
80 percent 1.774
100 percent 2.200

These calculated values show why the same decoding improvement can matter much more for generation-heavy search than for a verifier- or optimizer-dominated workload. They suggest measuring the time spent in generation, verification and updates before selecting where to invest systems effort. A change in the bottleneck changes the useful intervention, even when the attention architecture stays fixed.