Remark. What instruction tuning changes and leaves open [ftip-002I]
Remark. What instruction tuning changes and leaves open [ftip-002I]
Instruction tuning changes conditional likelihood under a declared demonstration distribution. It can teach response format, task interpretation, and behavior represented by the targets, but the training loss alone does not identify which of those effects caused an independent evaluation gain.
Ouyang et al. report supervised fine-tuning and its optimization settings in [ouyang2022training, Section 3.5 and Appendix C.1]. They do not define the exact mask and weighting convention displayed in Definition [ftip-002F]; those choices remain part of the declared objective.
The stage also leaves several questions open. A demonstration does not compare its target with alternatives, a teacher may transmit errors or hidden information, and low loss on the collected prompts need not imply transfer to a new task law. Preference acquisition introduces a different observation: which response a judge selected from a displayed pair.