AlphaEvolve: an AI-designed change actually used to train AI
AlphaEvolve: an AI-designed change actually used to train AI
Primary Novikov et al., “AlphaEvolve: A coding agent for scientific and algorithmic discovery” (Google DeepMind, June 2025), especially §§2, 3.3.2; authors’ May 2025 explanation. This is an observed engineering feedback path relevant to successor validation, not a general guide to coding agents; RSI synthesis should not say AI-designed changes have never entered model training.
Google researchers specify a tractable kernel-optimization problem, provide an initial annotated program, evaluation code and real TPU access; the Gemini-powered evolutionary agent proposes code changes. For a matrix-multiplication kernel used in Gemini training, the authors collect real input shapes, optimize on half and evaluate on the other half. They report a 23% average speedup of that kernel versus an expert-designed tiling heuristic and about 1% shorter overall Gemini training time. They explicitly say the heuristic was deployed in production; a separate scheduling heuristic recovered an average 0.7% of fleet compute after deployment, checked against live operation rather than simulator alone. For the kernel, preserving the same operation constrains correctness by construction. This is a genuine loop from an existing model to a narrower improvement in its own training machinery, not merely a benchmark win.
The scope boundary: humans supply the goal, scorable evaluators, hardware permission and production release path. Speeding a kernel does not show improvement in the successor model’s quality at fixed training compute, much less a sustained rise in quality-adjusted AI R&D productivity across generations. The 23% is for the tuned kernel; the end-to-end figure is 1%. The company reports those numbers; its paper does not give per-candidate all-in agent inference/TPU costs, a matched human/hybrid exploration budget, independent production logs or multiple successor generations. The training-shape split guards against fitting only seen shapes but does not measure research agenda selection. This is a strong counterexample to a claim that compute infrastructure and human engineers prevent all self-improvement, not evidence that they cease to be constraints.
Next test: obtain an audited series of promoted AI-origin improvements showing their direct effect on training cost and quality, failed candidates and all resource/approval costs, with matched human alternatives and how the next model’s research productivity changes. Keep the engineering feedback lane separate from short-horizon experiment judgment and alignment post-training.