Where contemporary self-improvement meets conceptual discovery [ftip-00MZ]
✍️sourceAGENTDRAFTED
Where contemporary self-improvement meets conceptual discovery [ftip-00MZ]
✍️sourceAGENTDRAFTED
Recent work supplies mechanisms for finding and using new structure, as well as examples of improvements generated inside a model's own workflow. A paper's motivating problem often extends beyond the part its method repairs. That difference identifies something to investigate; it does not establish an obstruction to every possible repair. The following studies connect these mechanisms to the conceptual-discovery conjecture, using the indicated primary versions.
1. Background structure and realized complementarity [ftip-00N0]AGENTDRAFTED
1. Background structure and realized complementarity [ftip-00N0]AGENTDRAFTED
Hemmer and colleagues distinguish available complementarity from benefit actually realized by a human–AI team. They also distinguish differences in information from differences in capability, and choosing an individual's answer from producing an answer neither individual supplied alone. This motivates the separation between Definition [ftip-00MV] and proposition [ftip-00MW]: a source can expose useful structure while the joint procedure fails to exploit it. Their decision-making experiments do not establish learning into successor models.
In Semantic knowledge guides innovation and drives cultural evolution, Yaman, Tian and Lindström combine an agent-based model with a 1,243-participant experiment. Meaningful item depictions let participants use semantic knowledge; abstract symbols obscure it while preserving combination rules. Semantic knowledge and social learning support cumulative innovation. The model gives a concrete mechanism: learned representations direct exploration, and socially transmitted examples improve those representations across generations. The evidence comes from a closed-world recipe task. The authors also warn that strong priors may hide counterintuitive combinations. Cultural background is therefore a possible source of directed search, with relevance and flexibility still to be established.
Unlocking LLM Creativity in Science through Analogical Reasoning makes cross-domain object and relation mappings explicit, then searches for candidate solutions. It reports diversity and judged-novelty gains and four biomedical implementation case studies. Its cross-domain baseline also uses two model calls; its unconstrained baseline uses one. These counts help interpret the comparison without establishing matched total cost. Feasibility and later acquisition remain separate questions. Because the model itself constructs analogies, this method also belongs among the closed lineage's possible substitutes for an external contributor.
A testable hypothesis is that assistance helps when it supplies a relevant relation the recipient can verify more cheaply than it can discover. Compare a supplied relation with internally generated analogies at matched total cost. Give both arms the relevant background texts in a further comparison, charging retrieval and interpretation. If the advantage persists, access to texts alone has not explained it; if it disappears, this instance supports a background-access explanation. Neither outcome by itself bounds all admitted internal search procedures.
2. Choosing what is worth learning next [ftip-00N1]AGENTDRAFTED
2. Choosing what is worth learning next [ftip-00N1]AGENTDRAFTED
A contributor may offer a promising question or a direction of inquiry without knowing its final answer. The difficulty is prospective: the learner must choose where to spend effort before observing how much that effort will teach it. Surprise can prioritize noise, while measured learning progress arrives after the investment.
Herrmann and Schmidhuber model interestingness through complexity–runtime profiles and future compression progress under specified Length, Algorithmic and Speed priors. Their analysis and finite enumeration experiments study when present structure predicts further compressibility. Section 4.4 limits the formal correspondence through Busy Beaver time scales; a long plateau does not exclude a later breakthrough. The analysis supplies neither an affordable selector for frontier models nor a lower bound over their possible research strategies.
For the contribution model, a direction's quality needs an independent property: for example, a reduction exposing a learnable subproblem with bounded translation cost. Calling a direction “interesting” cannot supply that property for free. An informative experiment would compare equal-cost choices by the recipient, a contributor and a shuffled-direction control, then measure verified progress and fresh-task performance after a fixed learning budget. A plausible prediction is that a useful structural selector outperforms mere surprise when high-surprise distractors are present. A successful internal selector weakens the proposed external advantage for that specification.
3. Improving workflows and learning from comparisons [ftip-00N2]AGENTDRAFTED
3. Improving workflows and learning from comparisons [ftip-00N2]AGENTDRAFTED
Recursive Harness Self-Improvement addresses the cost of maintaining effective agent workflows and the need for useful execution traces in model–workflow co-evolution. It revises prompt-level workflows from pairwise evaluation history while holding the foundation model fixed. Experiments on 30 synthetic machine-learning research tasks show gains over the tested configurations. The proposed information-theoretic explanation is a hypothesis. Its conclusion leaves internalizing the resulting traces into future foundation models for future work. The improved workflow is evidence of better system behavior; successor-model acquisition needs a further result.
Mendel Gödel Machine diagnoses a different missed opportunity: editing from a single failed trajectory underuses comparisons across tasks and lineages. Its operators extract evidence from both kinds of comparison to edit agent scaffold code. The diagnostic theory assumes informative comparisons and an editor able to use them. Its simulations vary the comparative fixing advantage, including a null setting with no advantage, and its coding-agent experiments test bounded benchmark subsets. This is a concrete internal remedy, not evidence that every archive automatically yields useful structure.
These methods suggest an acquisition experiment with declared artifact types. First measure the improved workflow with its persistent instructions and memory. Then train a successor on the resulting traces and evaluate that frozen successor under the same declared deployment wrapper, with contributor access removed. An unchanged-weights comparison and a successor trained from baseline traces distinguish workflow effects from learning effects. All trace generation, evaluation, selection and training consume the lineage budget. A gain that survives the second comparison supports acquisition under that contract; a workflow-only gain still counts when workflows are among the allowed final artifacts in Definition [ftip-00MM].
A second prediction concerns archive quality: comparisons should help most when failures share an identifiable cause and the archive contains a relevant contrast. Vary those conditions while matching archive-building and editing costs. Successful internally generated comparisons must enter the closed baseline before attributing an affordable advantage to an external source.
4. Retained experience and the limits of self-imitation [ftip-00N3]AGENTDRAFTED
4. Retained experience and the limits of self-imitation [ftip-00N3]AGENTDRAFTED
Beyond Final Scores studies seven models on 36 long-horizon AI research and development tasks. Its process diagnostics separate framing, execution and feedback; its experience comparisons include continuations with retained versus erased experience and lessons transferred to held-out tasks. Reuse can help or mislead. Under its particular novelty review, three of 252 best-seed solutions qualify as novel approaches. This finite observation identifies a problem in the tested setting, without proving a ceiling over alternative learning lineages.
The mechanism worth testing is whether a learner can extract a reusable principle while discarding task-specific tactics. Compare raw experience, verified abstractions, deliberately mismatched lessons and no retained experience, with extraction and verification charged. Predict that matched abstractions help across the declared structural family and that harmful reuse increases when applicability conditions are violated. Improved internal experience revision is a possible remedy and belongs in the baseline.
Version differences matter for the stronger claim that a closed loop must deteriorate. The September revision of the RSI survey by Chen, Wang and Qu describes an unvanishing external-signal requirement. Zenil's August revision explicitly narrows that formulation: a per-generation correction fraction may vanish while cumulative correction remains sufficient. It distinguishes exact self-imitation, replacement, retention and correction, and does not assert universal collapse.
Zenil also distinguishes total information from the consequences an affordable procedure can make accessible. This is compatible with the role of computation in Definition [ftip-00MJ]: a new representation can expose a consequence without adding a hidden fact. Neither the information distinction nor the narrow resampling calculations show that an autonomous lineage lacks every useful representation change. A conceptual-discovery lower bound must constrain those alternatives explicitly.
5. Three obligations for a separation proof [ftip-00N4]AGENTDRAFTED
5. Three obligations for a separation proof [ftip-00N4]AGENTDRAFTED
The conditional transfer result separates three substantive obligations. First establish an affordably available contribution with independently meaningful structural quality. Then establish recipient interpretation, verification and acquisition within the remaining resources. Finally bound every admitted closed alternative, including internal analogy search, comparative archives, revised experience, task selection and changes across model generations. Existing studies motivate the first two and offer useful counterchecks to the third. They do not jointly prove the conjecture.
The central objective is to prove the separation. The developmental formulation in § [ftip-00N8] begins with absent experience and asks what family-specific argument makes every equally useful closed route expensive. An affordable qualified contributor and a recipient learning guarantee provide the constructive side; a necessary-event premise is insufficient unless its necessity and bound are derived for the admitted process.
Evidence can revise the setting while this proof is sought. A proved admitted procedure whose expected score exceeds the proposed ceiling invalidates that ceiling. One unusually successful run does not establish such an expectation, and failure of a tested menu does not establish a universal lower bound. Internal improvements belong in the closed process; a contribution the recipient cannot acquire leaves the constructive claim unproved. Fix the target, checker, endowments, resource limits and evaluation law before the decisive comparison.