Alignment gain and preference information [ftip-008E]

Two finite models ask different questions about post-training. The first assumes an exact KL-regularized optimizer and characterizes the reward gain that its exponential tilt can produce. The second asks what ordinal comparisons reveal when post-training may reroute a fixed set of response circuits.

Neither source model is a general account of neural post-training. The first assumes exact optimization of a declared scalar reward. The second is a stylized routing model with a fixed circuit set. Extending the KL identity to an approximate optimizer requires further analysis. Extending the ordinal routing conclusion to a changing circuit set requires additional assumptions on that change.