Remark. Post-training with intermediate observations [ftip-004X]

A proposed interactive extension of the response-level RLVR cycle in § [ftip-0042] uses the environment of § [ftip-001W] and the action boundary of Definition [ftip-004P]. It retains intermediate observations and environment state instead of treating a whole response as one indivisible action.

This change introduces new questions about partial observability, long-horizon credit, environment versioning, recovery after interruption, and the cost of external execution. It does not assert that multi-turn training is uniformly better than response-level training.