Comparing controllers at fixed worker capability [ftip-00MD]
Comparing controllers at fixed worker capability [ftip-00MD]
To isolate coordination, fix the worker checkpoints, tool interfaces, prompt templates, evaluator, task law, and hard resource caps before varying the controller. Useful comparisons include a fixed serial policy, a static worker portfolio, a one-step value-of-information policy, and an adaptive policy that retains the full permitted history. Their allowed observations must agree. A solver that sees hidden outcomes has a different information interface; its optimum supplies an upper bound only through a justified relaxation containing the compared policies.
Use the controls of Definition [ftip-00JV] with controller logic declared as the varying coordinate. If controller-specific tuning is permitted, apply Definition [ftip-00JW]'s equal tuning data, selection, stopping, and cost rules. Compare suprema only within the intervention classes of Definition [ftip-00JX]. A separate factorial comparison can vary the checkpoint and controller independently, distinguishing learned capability, coordination, and their interaction.
Charge proposal and controller work, all parallel worker work, tool and solver use, communication, verification, failures, and retained-state updates using Definition [ftip-005H]. Report elapsed time alongside these costs: a shorter run with more workers may consume more total computation. Apply Definition [ftip-00CJ]'s admission rule before execution and record any contract violation separately from the quality of the submitted artifact.
Finite environments with rational transition laws can expose three distinct mechanisms: complementary observations, correlated worker failures, and delayed verification or newly revealed dependencies. Keep the latent state, observation rules, legal actions, and costs explicit. The first family includes Example [ftip-00M6]; the erased-bit construction in Example [ftip-00M7] tests whether a proposed summary loses decisive information. Where the exact model applies, compare attainable policy values with Remark [ftip-00MB]'s finite optimum and Theorem [ftip-00MA]'s upper certificate.
The principal quantity is the justified interval \([L_{\mathbf B},U_{\mathbf B}]\) for the restricted frontier. A smaller gap can result from a better feasible policy, a sharper upper certificate, or a more informative valid representation; distinguish these causes. In a live system, mean held-out transition accuracy does not justify a uniform model-error radius. Empirical quality and cost remain useful even when an upper certificate is unavailable.
Use evaluation instances separate from tuning, and declare task draws, repetitions, resource checkpoints, and analysis rules before comparison. Repeated runs share an instance and are not automatically independent tasks; account for this grouping in the analysis. Report independently verified quality, feasibility failures, time to the first valid artifact, and charged cost at matched quality. A valid confidence sequence can support repeated inspection under its statistical assumptions; see Howard et al.. It does not by itself justify selecting a new policy on reused evaluation outcomes.
These controls yield testable predictions. Complementary information can favor adaptation over one-step stopping; redundant workers can erase a portfolio's gain; delayed verification can make an early apparent success expensive to repair. An empirical claim of better coordination fails if its advantage disappears after matching access and charging the controller's own work. A claim of little remaining potential additionally requires the upper certificate, not merely a plateau among tested policies.