Remark. What instruction tuning changes and leaves open [ftip-002I]
AGENTDRAFTED
Instruction tuning changes conditional likelihood under a declared
demonstration distribution. It can teach response
format, task interpretation, and behavior represented by the targets, but the
training loss alone does not identify which of those effects caused an
independent evaluation gain.
Ouyang et al. report supervised fine-tuning and its optimization
settings in [ouyang2022training, Section 3.5 and Appendix C.1]. They do
not define the exact mask and weighting convention displayed in
Definition [ftip-002F]; those choices remain part of the declared objective.
The stage also leaves several questions open. A demonstration does not
compare its target with alternatives, a teacher may transmit errors or hidden
information, and low loss on the collected prompts need not imply transfer to
a new task law. Preference acquisition introduces a different observation:
which response a judge selected from a displayed pair.