Definition. Preference data law [ftip-002O]

A preference data law \(\mathcal D_{\mathrm {pref}}\) is a probability law on pairwise comparison observations \(o=(x,y^+,y^-,j)\) from Definition [ftip-002L]. It includes the randomness of prompt selection, response generation, pair selection, judge selection, and the judge's reported label.

A finite training multiset \(D_{\mathrm {pref}}=(o_1,\ldots ,o_N)\) is sampled or adaptively collected under that law and its acquisition history. When queries are adaptive, the records need not be independent or identically distributed.