Reinforcement-learning foundations [ftip-002W]

Reinforcement learning updates a policy from rewards attached to sampled behavior. Before considering a particular alignment method, this section fixes the response-level episode, return, value, advantage, policy roles, and likelihood ratios used by policy-gradient objectives.