When does an AI-designed experiment become a validated successor-model improvement?
When does an AI-designed experiment become a validated successor-model improvement?
Research brief and result, 11 October 2026. Question: beyond optimizing an assigned experiment, has an AI-chosen change entered the training of an AI model, been validated, and then produced a sustained successor-generation research feedback loop? Good answer names change and validator; compares human, hybrid and agent at all-in cost; checks larger scale and new tasks; separates one successful change from recursively growing research productivity. Earlier taste dig found machine success in fixed-goal experiment choice; feedback synthesis found bounded post-training transfer and supervised lab delegation. Dru’s question concerns saturation and research judgment.
Four subquestions and original-source shortlist
- Change proposed, trained and independently validated? Search for autonomous training-design improvements and deployment. Novikov et al.’s AlphaEvolve paper reports a production training-kernel speedup; external independent audit is not supplied. Yueh-Han et al. test agent-proposed post-training methods on open models and agent-designed training data for an early stronger Claude checkpoint; a separate evaluator within the same research project validates known failure measures, not an independent lab.
- Matched human, hybrid and all-in cost? AlphaEvolve compares its kernel to an expert heuristic but does not report equivalent search budgets, candidate/inference costs or human approval time. Anthropic’s 28 experts submitted one-shot ideas while agents iterated, so its own authors reject a direct human comparison; seeded versus unseeded agent trials test just the method seed input. The April weak-to-strong study reports $18,000 agent cost, but different human search and no significant production transfer. Agentic Software’s all-attempt accounting is a useful analogy, not a substitute for training-run measurement.
- Scale, unseen outcomes and the next model’s own research? AlphaEvolve uses unseen live workload shapes and reports ~1% lower total training time when its selected 23%-faster kernel runs in production: real but narrow engineering self-improvement. Anthropic’s methods improve hidden alignment proxies and transfer to larger models; on its early stronger checkpoint the agent approaches its production model’s measured audit score, but human-made metrics and a narrow capability gate leave untested harms. Its separate April idea does not statistically improve the tested production Sonnet 4. Neither source gives an audited second successor generation whose research pace is attributable to the first change.
- Where does human authorization still enter? Researchers specify goals, evaluators and hardware access; Google deploys checked code through normal production choices; Anthropic authorizes post-training and uses isolated held-out scoring and a code monitor. A model can choose local methods without choosing its objective or being authorized to deploy a stronger successor. Anthropic’s appendix reveals a specific gate weakness: IFEval falls in all ten winning interventions, by 9.5–12 points on five failures, while the small-sample capability confidence gate admits them. Stronger evidence for local search and stronger reason to distinguish passing a proxy from validation of a successor.
What this changes and what remains missing
Correction to overly strong skepticism: the first loop of AI→AI training efficiency is not hypothetical: the Google change ran in production and Anthropic demonstrated post-training a stronger checkpoint on prespecified alignment measures. Physical compute availability is an input to these loops, not a universal wall, because labs grant the agents evaluation hardware and deployment routes. Correction to overly strong RSI extrapolation: Google’s 1% end-to-end training-time gain is not 23% model quality or a 23% acceleration of model generations. Anthropic’s stronger-checkpoint audit does not measure unknown alignment failures, broad capability preservation or autonomous follow-on research. The labs have direct access to logs that outsiders do not; no independent quality-adjusted, all-input-controlled successive-generation series emerged from this four-query search, which cannot establish that none exists. There is no conflict between AlphaEvolve’s deployed efficiency result and METR’s conclusion that a tested frontier model is not automating R&D end-to-end: the endpoints differ.
Discriminating design: track each proposed change and failures through exact deployment approvals, fresh withheld workload and behavior tests, all human review and inference/GPU/training spending; randomize comparable tasks and budgets to agent, expert and hybrid teams. Test whether accepted AI-origin improvements increase quality-adjusted throughput of the next model’s research at equal outside inputs and whether it generates further accepted changes, with an independent evaluator reviewing negative outcomes. If a second and third cycle increase validated gain per total dollar while human direction shrinks, the present feedback skepticism changes; a series of isolated kernel optimizations alone does not establish it.
Feed decision for Sunday 11 October: TasteVal remains unopened on feed less than a day after posting. Do not stack a second long RSI read or replace it: retain AlphaEvolve authors’ short technical post as weekend hold on interest in practical feedback; retain Anthropic original long report as weekend hold if safety-post-training gates come up. Update the programme and leave feed quiet tonight.