Finite consequences and counterexamples [ftip-00JE]
✍️sourceAGENTDRAFTED
Finite consequences and counterexamples [ftip-00JE]
✍️sourceAGENTDRAFTED
Finite sampling, reward information, and deterministic execution give conditional conclusions about post-training. Their counterexamples show why the task law, feasible set, and available information cannot be omitted.
This synthesis brings together the finite results used in the evaluation analysis. The conceptual-discovery chapter asks whether such arguments can constrain a complete learning lineage, including changes to representations, curricula and research procedures.
1. Probability, information, and execution [ftip-00I1]AGENTDRAFTED
1. Probability, information, and execution [ftip-00I1]AGENTDRAFTED
Discovery probabilities, support, proxy error, feedback resolution, and replay each constrain a different part of a comparison. Capability remains conditional behavior under a declared intervention and evaluation law, rather than a scalar latent property.
1.1. Finite probability and execution results [ftip-00I2]AGENTDRAFTED
1.1. Finite probability and execution results [ftip-00I2]AGENTDRAFTED
The conclusions depend on their declared probability laws, finite index sets, positivity conditions, and execution inputs. Changing one of these assumptions can change the result even when the observed score is unchanged.
Convention 1.1.1. Domains of the finite results [ftip-00I5]AGENTDRAFTED
Convention 1.1.1. Domains of the finite results [ftip-00I5]AGENTDRAFTED
Each result concerns its declared carriers \(X_1,\ldots ,X_n\), with fixed equality and order conventions. Finiteness, probability laws, and positivity conditions are hypotheses of the result; they cannot be inferred from a finite observation alone.
Theorem 1.1.2. Independent discovery probability [ftip-00I7]AGENTDRAFTED
Theorem 1.1.2. Independent discovery probability [ftip-00I7]AGENTDRAFTED
For \(B\in \mathbb N\) independent discovery events with the probabilities declared in Convention [ftip-0077], the identity in Theorem [ftip-0078] gives \(\Pr (D_B)=1-\prod _{i=1}^{B}(1-p_i)\).
Theorem 1.1.3. Discovery budget for a failure threshold [ftip-00I8]AGENTDRAFTED
Theorem 1.1.3. Discovery budget for a failure threshold [ftip-00I8]AGENTDRAFTED
Under \(0<p<1\), \(0<\delta <1\), and an integer budget \(B\geq 1\), the failure constraint is equivalent to the following bound.
\[B\geq \left \lceil \frac {\log \delta }{\log (1-p)}\right \rceil .\]The result is proved in Corollary [ftip-0079].
Theorem 1.1.4. Support under finite exponential tilting [ftip-00I9]AGENTDRAFTED
Theorem 1.1.4. Support under finite exponential tilting [ftip-00I9]AGENTDRAFTED
Under the finite-carrier and positive-temperature hypotheses of Theorem [ftip-007C], normalized exponential tilting preserves the base law's positive support.
Theorem 1.1.5. Uniform proxy error and objective regret [ftip-00IA]AGENTDRAFTED
Theorem 1.1.5. Uniform proxy error and objective regret [ftip-00IA]AGENTDRAFTED
For a finite common feasible set and a common regularizer, the uniform proxy error hypothesis in Definition [ftip-007G] yields the two-epsilon objective regret bound proved in Theorem [ftip-007H]; the conclusion concerns the regularized objective named there.
Theorem 1.1.6. Utility under a change of evaluation law [ftip-00IB]AGENTDRAFTED
Theorem 1.1.6. Utility under a change of evaluation law [ftip-00IB]AGENTDRAFTED
If the evaluation laws and finite response space satisfy the total-variation hypotheses of Theorem [ftip-00GM], then the expectation difference is bounded by the stated sup-norm times total variation.
Theorem 1.1.7. Response classes under feedback refinement [ftip-00IC]AGENTDRAFTED
Theorem 1.1.7. Response classes under feedback refinement [ftip-00IC]AGENTDRAFTED
For finite response set \(\mathcal Y\) and deterministic rewards, a reward channel that refines the equality partition of another channel separates at least as many response classes, as proved in Lemma [ftip-00AG].
Theorem 1.1.8. Equal execution inputs give equal replay traces [ftip-00IJ]AGENTDRAFTED
Theorem 1.1.8. Equal execution inputs give equal replay traces [ftip-00IJ]AGENTDRAFTED
If the execution map is deterministic in all coordinates fixed by Definition [ftip-00HF], then equal input, revision, environment, and seed records produce equal finite traces, as proved in Theorem [ftip-00HG].
Theorem 1.1.9. Observational equivalence under a common update kernel [ftip-00IK]AGENTDRAFTED
Theorem 1.1.9. Observational equivalence under a common update kernel [ftip-00IK]AGENTDRAFTED
For the finite transcript and world kernels declared in Definition [ftip-008C]--Definition [ftip-008D], observationally equivalent worlds induce the same output law for any common randomized post-training kernel, as proved in Theorem [ftip-007S].
Theorem 1.1.10. Conditions for an admissible commit decision [ftip-00IL]AGENTDRAFTED
Theorem 1.1.10. Conditions for an admissible commit decision [ftip-00IL]AGENTDRAFTED
Under the typed audit record of Definition [ftip-00HL], a commit is admissible exactly when all required checks pass, the digest matches, and the recorded decision is \(commit\), as proved in Theorem [ftip-00HM].
Theorem 1.1.11. Admission of a finite audit plan [ftip-00IM]AGENTDRAFTED
Theorem 1.1.11. Admission of a finite audit plan [ftip-00IM]AGENTDRAFTED
A finite audit plan with nonnegative stage costs is admissible exactly when its declared additive cost is within budget; adding a positive stage beyond slack is inadmissible, by Theorem [ftip-00HX].
1.2. Counterexamples and restricted assumptions [ftip-00ID]AGENTDRAFTED
1.2. Counterexamples and restricted assumptions [ftip-00ID]AGENTDRAFTED
Finite constructions can separate average success from coverage, proxy accuracy from evaluation quality, and observed gradients from the hypotheses needed to bound them. Each separation identifies an assumption that a stronger conclusion would require.
Remark 1.2.1. What a finite counterexample establishes [ftip-00IE]AGENTDRAFTED
Remark 1.2.1. What a finite counterexample establishes [ftip-00IE]AGENTDRAFTED
A finite assignment satisfying a claim's hypotheses and violating its conclusion disproves that universal claim. It does not estimate how often the failure occurs under another distribution or establish a general failure rate.
Remark 1.2.2. Changing the model or evaluation changes the claim [ftip-00IF]AGENTDRAFTED
Remark 1.2.2. Changing the model or evaluation changes the claim [ftip-00IF]AGENTDRAFTED
A conclusion about fixed carriers, policies, feedback, budgets, and evaluation laws need not survive a change to those objects. Finite zero hits do not imply zero support, empirical benchmark movement does not imply acquisition, and an exact-optimizer result need not hold for an approximate optimizer without an additional error bound.
Example 1.2.3. Equal averages and zero hits do not identify coverage [ftip-00IG]AGENTDRAFTED
Example 1.2.3. Equal averages and zero hits do not identify coverage [ftip-00IG]AGENTDRAFTED
The two-task examples in Example [ftip-007E] and Example [ftip-00A3] show that equal one-shot averages or zero observed hits can coexist with different coverage or nonzero latent rates. These conclusions use their declared iid model and do not estimate rates outside it.
Example 1.2.4. Small training-law error can reverse evaluation selection [ftip-00IH]AGENTDRAFTED
Example 1.2.4. Small training-law error can reverse evaluation selection [ftip-00IH]AGENTDRAFTED
In the two-point construction of Example [ftip-007J], training-law average error is small while evaluation selection reverses. The missing hypothesis is a uniform bound on the evaluated feasible set; an informal promise of similar distributions does not supply it.
Remark 1.2.5. Untied heads, shared gradients, and checkpoint forecasts [ftip-00II]AGENTDRAFTED
Remark 1.2.5. Untied heads, shared gradients, and checkpoint forecasts [ftip-00II]AGENTDRAFTED
For DGG, the untied head identity and its finite-batch Jensen bound in Theorem [ftip-007Z] and Theorem [ftip-0081] differ from an intermediate-occurrence bound. A shared intermediate weight requires the sum over all occurrences in Theorem [ftip-007Z]; the source's local hypotheses do not establish its claimed total-gradient bound in [miao2026when, Section 4.2.1, Lemma 1, Theorem 1, and Appendix A.3]. The cancellation example Example [ftip-0084] also separates the active clipping interval from the raw-surrogate extension.
NExt's empirical checkpoint forecasts in [chen2026lowrank, Sections 3.2, 4.1--4.3, and 5.1] do not establish a universal subspace-collapse, capability, or speed theorem. The spectral and evaluation quantities in Definition [ftip-0086] through Remark [ftip-008B] measure different trajectory properties.
1.3. Mathematical notation [ftip-00IN]AGENTDRAFTED
1.3. Mathematical notation [ftip-00IN]AGENTDRAFTED
Prompt and response spaces, probability laws, policies, rewards, evaluation scores, and execution traces are distinct mathematical objects.
Convention 1.3.1. Prompt spaces, response spaces, and laws [ftip-00IO]AGENTDRAFTED
Convention 1.3.1. Prompt spaces, response spaces, and laws [ftip-00IO]AGENTDRAFTED
Use \(\mathcal X\) for prompts, \(\mathcal Y\) for finite responses, \(\Omega \) for sample space, \(\mathsf P\) for a declared protocol, and \(\mu \) for a prompt law. A probability law is written \(\Pr \) only after its sample space is named.
Convention 1.3.2. Policies and reference policies [ftip-00IP]AGENTDRAFTED
Convention 1.3.2. Policies and reference policies [ftip-00IP]AGENTDRAFTED
Use \(\pi \) for a policy, \(\pi _{\rm base}\) for a reference policy, and \(\pi _\theta \) for a parameterized policy. A policy is always typed as a kernel from the declared prompt carrier to the declared response carrier.
Convention 1.3.3. Rewards and evaluation utilities [ftip-00IQ]AGENTDRAFTED
Convention 1.3.3. Rewards and evaluation utilities [ftip-00IQ]AGENTDRAFTED
Use \(r:\mathcal X\times \mathcal Y\to \mathbb R\) for a deterministic reward and \(u\) for an evaluation utility. Proxy, verifier, and process rewards retain their local subscripts; no bare symbol silently changes meaning.
Convention 1.3.4. Model artifacts and evaluation scores [ftip-00IR]AGENTDRAFTED
Convention 1.3.4. Model artifacts and evaluation scores [ftip-00IR]AGENTDRAFTED
Use \(M\) for a model artifact, \(\mathsf E\) for an evaluation protocol, and \(J_{\mathsf E}(M)\) for its declared score. A comparison must state the invariant evaluation law before subtracting two scores.
Convention 1.3.5. Traces, states, lineages, and costs [ftip-00IS]AGENTDRAFTED
Convention 1.3.5. Traces, states, lineages, and costs [ftip-00IS]AGENTDRAFTED
Use \(z_{0:n}\) for a finite trace, \(s_t\) for a harness state, \(\ell \) for a lineage, and \(c\) for an execution cost. Event logs and summaries are distinct carriers even when one is computed from the other.