Remark. Evaluation gain under a fixed task law [ftip-005G]

The policy-evaluation expectation of Section 3.5 of Reinforcement learning: An introduction[sutton2018reinforcement] specializes to an independently fixed task law and inference procedure. The gain is a paired contrast with the same base artifact and evaluator.

A positive value shows a change under this evaluation. It does not alone show that the trained policy acquired a capability under the criterion of Definition [ftip-0007].