Improving workflows and learning from comparisons [ftip-00N2]

Recursive Harness Self-Improvement addresses the cost of maintaining effective agent workflows and the need for useful execution traces in model–workflow co-evolution. It revises prompt-level workflows from pairwise evaluation history while holding the foundation model fixed. Experiments on 30 synthetic machine-learning research tasks show gains over the tested configurations. The proposed information-theoretic explanation is a hypothesis. Its conclusion leaves internalizing the resulting traces into future foundation models for future work. The improved workflow is evidence of better system behavior; successor-model acquisition needs a further result.

Mendel Gödel Machine diagnoses a different missed opportunity: editing from a single failed trajectory underuses comparisons across tasks and lineages. Its operators extract evidence from both kinds of comparison to edit agent scaffold code. The diagnostic theory assumes informative comparisons and an editor able to use them. Its simulations vary the comparative fixing advantage, including a null setting with no advantage, and its coding-agent experiments test bounded benchmark subsets. This is a concrete internal remedy, not evidence that every archive automatically yields useful structure.

These methods suggest an acquisition experiment with declared artifact types. First measure the improved workflow with its persistent instructions and memory. Then train a successor on the resulting traces and evaluate that frozen successor under the same declared deployment wrapper, with contributor access removed. An unchanged-weights comparison and a successor trained from baseline traces distinguish workflow effects from learning effects. All trace generation, evaluation, selection and training consume the lineage budget. A gain that survives the second comparison supports acquisition under that contract; a workflow-only gain still counts when workflows are among the allowed final artifacts in Definition [ftip-00MM].

A second prediction concerns archive quality: comparisons should help most when failures share an identifiable cause and the archive contains a relevant contrast. Vary those conditions while matching archive-building and editing costs. Successful internally generated comparisons must enter the closed baseline before attributing an affordable advantage to an external source.