Post-training, alignment, and feedback [ftip-00JA]

Post-training procedures differ in the information they receive and the updates they permit. Demonstrations, preferences, rewards, verifiers, and environment interactions impose different conditions on the resulting policy.

The chapter moves from demonstrations and preferences to reward-based updates, then to training through tool interaction and replay. The agent-state analysis complements these parameter-changing procedures by examining the contexts, memories and computation an agent retains around its models.