Example. Same model under two procedures and budgets [ftip-0004]

Holding the model fixed, the comparison isolates how inference procedure and budget alter measured success.

Assume independent draws and a declared perfect checker that recognizes a correct candidate when one appears. The one-draw procedure succeeds with \(J_1=p\). A procedure allowed up to \(k\) draws succeeds whenever not all draws fail: \[ J_k=1-(1-p)^k. \] Both procedures query the same conditional distribution.

Policy performance under a specified interaction and evaluation procedure is standard in [sutton2018reinforcement, Section 3.5]. The formula demonstrates procedural dependence. Independence and a perfect checker are declared toy assumptions, not properties attributed to a frontier language model.