Costs, measurements and outcomes that would change the claim [ftip-00NV]
Costs, measurements and outcomes that would change the claim [ftip-00NV]
A campaign's resource record includes model inference, training, mathematical experiments, classical computation, retrieval, contribution production, communication, certificate generation and verification. It also includes failed attempts, discarded checkpoints and validation used to select a method. Model calls alone do not equalize systems using different models, context lengths or training procedures. Record accelerator time, CPU work, human time, money, memory and wall time separately before applying any declared scalar conversion.
Let \(D_A\) denote development work for an autonomous process and \(D_C\) the contributor's charged development. Let \(T\) include transmission, recipient adaptation and validation. For a common deployment budget, write \(E_A(N)\) and \(E_C(N)\) for work on \(N\) fresh instances, including certification. A specified accounting convention then compares
\[ \begin {aligned} W_A(N)&=D_A+E_A(N),\\ W_C(N)&=D_C+T+E_C(N). \end {aligned} \notag\]The expression includes any autonomous reconstruction of contributor preparation in \(D_A\). Report inherited preparation and resources treated as sunk separately for both sides. A marginal-session saving does not establish a lifetime saving. If resources remain a vector, compare each component or state a Pareto relation; do not add human hours to accelerator operations without a conversion rule.
At equal task distributions and accepted-success requirements, reuse can amortize a contribution. Under the additional approximation of constant per-instance costs \(e_A\) and \(e_C\), with \(e_A>e_C\), the contributed method has lower total work precisely when
\[ N(e_A-e_C)>D_C+T-D_A. \notag\]This is an accounting consequence, not a prediction that the inequality holds. If the contributed method has lower success, comparisons must account for the additional work required to reach the same criterion. If its per-instance work is no better, amortization alone cannot erase a larger initial cost. Report failures at a cap as censored attempts; do not turn them into an infinite measured cost or a solved-instance ratio.
The primary acquisition measurement is the fraction of new instances with accepted certificates under the fixed deployment cap. Record it separately for the original distribution and extrapolation regime, along with certificate lengths, work and failures. Repeat complete development campaigns, rather than only decoding from one favorable learned library. Use common evaluation instances for paired comparisons and report variation across development seeds. Cost-to-threshold conclusions require uncertainty for both success and work and must name the procedures and budgets tested.
Additional measurements explain a result without replacing it: which verified lemmas are used in accepted proofs; whether a transformation method handles new presentations; whether the recipient still succeeds without contributor access; and whether a new recipient benefits from the same artifact. A proof differing textually from examples is not by itself evidence of conceptual novelty. Conversely, reuse of a short verified lemma can be a useful acquisition even when that lemma is familiar to mathematicians.
The literature motivates several testable predictions.
- Relevant verified lemmas should improve recipient performance more than similarly sized irrelevant examples, as suggested by the reuse studies in DreamProver and ProofEvolve. A missing difference would weaken the claimed library mechanism for this family.
- Challenging conjectures with counterexamples should expose the failure of incomplete invariants, including the middle-degree example in § [ftip-00NS]. If a simple algebraic solver already gives equivalent performance at lower cost, this mechanism supplies no cost advantage.
- A useful curriculum should improve frozen-recipient performance beyond direct exposure to its material at the same accounted resources. If it merely helps while the teacher remains present, the result concerns assisted deployment rather than retained capability.
- Amortized savings, when present, should depend on reuse count and the deployment regime. Failure on larger compositions would restrict transfer even if the original-size evaluation improves.
A positive pilot would establish a bounded acquisition result for the specified procedures. An autonomous reconstruction at comparable cost would weaken a proposed developmental advantage; a cheap classical solver would defeat an all-method barrier for this family. Neither result decides whether a different mathematical research family admits the separation in § [ftip-00MN]. Moving to that stronger claim requires a new family and an argument covering all equally useful methods, including implicit representations and future model generations.