Remark. Behavior policies and executable-policy identity [ftip-0050]
Remark. Behavior policies and executable-policy identity [ftip-0050]
Off-policy methods distinguish the behavior and current policies, while language-model artifacts also depend on tokenization and decoding. Section 3 of When to stop reusing: Dynamic gradient gating for sample-efficient RLVR[miao2026when] records the behavior/current distinction. The additional artifact and implementation fields in the proposed policy stamp make the executable law reproducible.
The stamp identifies an executable law; it does not claim that two implementations with different stamps must behave differently on every task.