Counterexamples and restricted assumptions [ftip-00ID]
✍️sourceAGENTDRAFTED
Counterexamples and restricted assumptions [ftip-00ID]
✍️sourceAGENTDRAFTED
Finite constructions can separate average success from coverage, proxy accuracy from evaluation quality, and observed gradients from the hypotheses needed to bound them. Each separation identifies an assumption that a stronger conclusion would require.
Remark 1. What a finite counterexample establishes [ftip-00IE]AGENTDRAFTED
Remark 1. What a finite counterexample establishes [ftip-00IE]AGENTDRAFTED
A finite assignment satisfying a claim's hypotheses and violating its conclusion disproves that universal claim. It does not estimate how often the failure occurs under another distribution or establish a general failure rate.
Remark 2. Changing the model or evaluation changes the claim [ftip-00IF]AGENTDRAFTED
Remark 2. Changing the model or evaluation changes the claim [ftip-00IF]AGENTDRAFTED
A conclusion about fixed carriers, policies, feedback, budgets, and evaluation laws need not survive a change to those objects. Finite zero hits do not imply zero support, empirical benchmark movement does not imply acquisition, and an exact-optimizer result need not hold for an approximate optimizer without an additional error bound.
Example 3. Equal averages and zero hits do not identify coverage [ftip-00IG]AGENTDRAFTED
Example 3. Equal averages and zero hits do not identify coverage [ftip-00IG]AGENTDRAFTED
The two-task examples in Example [ftip-007E] and Example [ftip-00A3] show that equal one-shot averages or zero observed hits can coexist with different coverage or nonzero latent rates. These conclusions use their declared iid model and do not estimate rates outside it.
Example 4. Small training-law error can reverse evaluation selection [ftip-00IH]AGENTDRAFTED
Example 4. Small training-law error can reverse evaluation selection [ftip-00IH]AGENTDRAFTED
In the two-point construction of Example [ftip-007J], training-law average error is small while evaluation selection reverses. The missing hypothesis is a uniform bound on the evaluated feasible set; an informal promise of similar distributions does not supply it.
Remark 5. Untied heads, shared gradients, and checkpoint forecasts [ftip-00II]AGENTDRAFTED
Remark 5. Untied heads, shared gradients, and checkpoint forecasts [ftip-00II]AGENTDRAFTED
For DGG, the untied head identity and its finite-batch Jensen bound in Theorem [ftip-007Z] and Theorem [ftip-0081] differ from an intermediate-occurrence bound. A shared intermediate weight requires the sum over all occurrences in Theorem [ftip-007Z]; the source's local hypotheses do not establish its claimed total-gradient bound in [miao2026when, Section 4.2.1, Lemma 1, Theorem 1, and Appendix A.3]. The cancellation example Example [ftip-0084] also separates the active clipping interval from the raw-surrogate extension.
NExt's empirical checkpoint forecasts in [chen2026lowrank, Sections 3.2, 4.1--4.3, and 5.1] do not establish a universal subspace-collapse, capability, or speed theorem. The spectral and evaluation quantities in Definition [ftip-0086] through Remark [ftip-008B] measure different trajectory properties.