Remark. Finite questions and necessary assumptions [ftip-0068]
Remark. Finite questions and necessary assumptions [ftip-0068]
Finite analysis raises three questions: how conditional success rates bound long-horizon success, how replay-ratio concentration controls update error, and how trajectory extrapolation with periodic recovery affects evaluation regret. A cost comparison also depends on rollout, update, environment, storage, and evaluation work.
Finite counterexamples show why these questions need additional assumptions. Finite policies can separate the entropy statistics above. Two verifiers can agree on all observed training records yet differ on an unobserved adversarial region. Two parameter paths can share early low-rank summaries and diverge after a new successful rollout. These examples mark where additional assumptions are unavoidable.
The DGG and NExt papers motivate reuse and trajectory-compression hypotheses but do not establish these general results. Bounds on update error and extrapolation regret require assumptions relating the monitored quantities to the corresponding outcomes.