Evaluation, evidence, and finite limits [ftip-00JB]

The training and agent mechanisms described earlier produce an artifact whose capability must be evaluated under a declared task law, inference procedure and resource account. This chapter fixes that comparison, studies what survives a change of evaluation, and examines discovery probabilities, support, proxy error and the information available through feedback.

Finite results and counterexamples establish consequences under specific probability laws and computational assumptions. They provide tools for the model-lineage question, where their assumptions must cover an evolving research and training process. The architecture analysis uses the same evaluation framework for comparisons that depend on a model implementation or optimizer.