Remark. Replay selection and current-policy annotations [ftip-0058]

The proposed replay interface separates a persistent pool from annotations computed for an update attempt. DGG in [miao2026when, Section 3; Section 5 and Algorithm 1] recomputes policy ratios and a gradient diagnostic while reusing a rollout batch. Those quantities depend on the current policy and therefore belong to the attempted update rather than the historical record.

The split also exposes the cost of selecting, scoring, and rejecting reused records.