Anthropic’s “When AI builds itself”: tasks accelerated, loop still open
Anthropic’s “When AI builds itself”: tasks accelerated, loop still open
Primary source: Anthropic Institute, “When AI builds itself”, checked 30 September 2026. Linked question: AI-assisted R&D versus recursive loop. This note concerns R&D feedback and safety, not advice on day-to-day coding workflows (Agentic Software).
What the lab claims
As of May 2026, over 80% of merged code was AI-authored; typical engineer merged ~8× as much code per day as in 2024, which Anthropic explicitly calls an overestimate of productivity. In a March poll of 130 research employees, median subjective output uplift with Mythos Preview versus no AI was ~4×; Anthropic expects the true gain to be lower. For a pre-defined small-model-training speedup task, it reports ~52× on the starting code with Mythos Preview in April (versus ~3× with Opus 4 in May 2025). Its 129 retrospective comparisons of researcher detours were selected because the human made a suboptimal choice, with another model serving as judge; the 64% favorable-to-Mythos figure is not a representative researcher-replacement test.
In an April 2026 weak-to-strong supervision experiment, nine agents achieved 0.97 performance-gap recovery over 800 cumulative agent-hours and about $18,000 versus two humans obtaining 0.23 in seven days on a small open model. Humans selected the problem, scoring and resources. Its strongest idea transferred to held-out math/coding datasets but did not statistically improve the tested production-scale Claude Sonnet 4. A separate 28 August automated-alignment study shows iterative improvement on ten specified, measured failure categories, including withheld benchmarks and larger test models, but still not general-purpose verified safe successors. The human comparison did not let the humans iterate on submissions.
The causal gap the author concedes
Anthropic says human problem choice, interpreting experiments, code review, experiment compute and eventually chips/power can become binding. A strong mini-loop within a given score is different from choosing an ungameable target, producing a frontier successor, validating the gain after training, and then repeating. Its author explicitly says fully autonomous RSI is neither present nor inevitable, and even if achieved would not remove material and institutional bottlenecks elsewhere.
Reading rule
Do not read a code-volume multiplier as a multiplier on capability progress, or a 97% small-model scoring result as 97% of a solved alignment problem. But neither is it credible to deny fast-growing capacity for well-scored experimental iteration; compare with METR’s September evaluation and cost-calibrated experiment.