Matched-compute comparison protocol [ftip-00JU]
✍️sourceAGENTDRAFTED
Matched-compute comparison protocol [ftip-00JU]
✍️sourceAGENTDRAFTED
Definition 1. Common-recipe comparison arm [ftip-00JV]AGENTDRAFTED
Definition 1. Common-recipe comparison arm [ftip-00JV]AGENTDRAFTED
A common-recipe arm fixes training data, optimizer family, schedule, post-training procedure and feedback, training and post-training budgets, inference budget, and evaluation protocol and draws. Interface-compatible parameter choices are declared in advance as part of each permitted architecture package. The comparison varies that package under the common recipe; any further deviation is recorded as a separate comparison coordinate. Its score difference measures performance under this recipe, not the best attainable result for either architecture.
Definition 2. Equal-tuning-budget comparison arm [ftip-00JW]AGENTDRAFTED
Definition 2. Equal-tuning-budget comparison arm [ftip-00JW]AGENTDRAFTED
An equal-tuning-budget arm permits architecture-specific tuning under a predeclared protocol that fixes the tuning data and its volume, trial count, selection rule, stopping rule, and tuning cost for both architectures. The cost uses the same declared accounting rule. The resulting comparison measures practical performance under this equal optimization effort; it does not determine best possible training or a representation-only limit.
Definition 3. Restricted-envelope comparison arm [ftip-00JX]AGENTDRAFTED
Definition 3. Restricted-envelope comparison arm [ftip-00JX]AGENTDRAFTED
A restricted-envelope arm declares an allowed intervention class \(\mathcal E_A\subseteq \mathfrak I_A\) for each architecture, specifying the permitted training, post-training, and inference procedures. The classes expose exclusions, search budgets, seeds, and stopping rules. Restrict the evaluation functional of Definition [ftip-00JJ] to \(\mathcal E_A\) and compare the resulting suprema under the same finite cost cap \(C\geq 0\), common evaluation law, and common declared cost accounting.
The measurable, integrable evaluation domain and extended-real conventions of Definition [ftip-00JJ] apply: an empty feasible class has value \(-\infty \), an unbounded-above feasible score set has value \(+\infty \), and a finite supremum need not be attained. The allowed class is part of the comparison, so its envelope is a restriction-dependent quantity, not an unrestricted architectural ceiling.
Definition 4. Model-instance comparison evidence record [ftip-00JY]AGENTDRAFTED
Definition 4. Model-instance comparison evidence record [ftip-00JY]AGENTDRAFTED
A model-instance record contains architecture/checkpoint identity, parameter count, data and post-training recipe, optimizer, precision, hardware, cache policy, training and inference cost vectors, task suite, and independent evaluation seed. A missing field makes an architecture attribution unknown.
Remark 5. Systems optimizations are measured interventions [ftip-00JZ]AGENTDRAFTED
Remark 5. Systems optimizations are measured interventions [ftip-00JZ]AGENTDRAFTED
FlashAttention, KV-cache layouts, quantization, batching, and kernels can change feasible cost. If they approximate, truncate, or alter precision of the mathematical computation, the record must mark that as a systems intervention rather than as a pure architecture comparison.
Theorem 6. Matched-protocol difference is an estimand [ftip-00K0]AGENTDRAFTED
Theorem 6. Matched-protocol difference is an estimand [ftip-00K0]AGENTDRAFTED
Given \(n\geq 1\) evaluation draws, let \(\widehat V_A(C)\) be the sample mean of finite real evaluation scores for architecture \(A\) at a finite cost cap \(C\geq 0\). Under a fixed evaluation law and common-recipe arm, the finite difference \(\widehat V_A(C)-\widehat V_B(C)\) is a well-defined empirical estimand.
It does not identify a causal architecture effect when data, tuning, or systems coordinates differ. The conclusion follows directly from the declared record.
Remark 7. Conditions for applying a toy model [ftip-00K1]AGENTDRAFTED
Remark 7. Conditions for applying a toy model [ftip-00K1]AGENTDRAFTED
A toy theorem applies to a concrete model only if an explicit map from its state and interface to the model preserves the theorem's assumptions, including its cost accounting. Without such a map, the conclusion concerns only the abstract model; its validity for the concrete system is unknown.