Notation. Prompts, responses, demonstrations, and loss masks [ftip-002A]
Notation. Prompts, responses, demonstrations, and loss masks [ftip-002A]
Let \(\mathcal X_{\mathrm {pr}}\) be the prompt space. For each prompt \(x\in \mathcal X_{\mathrm {pr}}\), let \(\mathcal Y(x)\subseteq \mathcal V^*\) be the set of finite token responses permitted by the training format. We write \(y=(y_1,\ldots ,y_{|y|})\in \mathcal Y(x)\) and \(y_{<t}=(y_1,\ldots ,y_{t-1})\).
A demonstration is denoted \(z=(x,y,\omega )\), where \(\omega \) is its provenance record. A finite multiset of demonstrations is denoted \(D_{\mathrm {sft}}\). For each response position, a loss mask \(m_t(z)\in \{0,1\}\) records whether the token contributes to the supervised objective, and a nonnegative weight \(w_t(z)\) records its declared normalization. Prompt tokens are conditioning context and lie outside the response-position domains of \(m_t(z)\) and \(w_t(z)\).