AI-assisted AI research versus a self-sustaining recursive loop
AI-assisted AI research versus a self-sustaining recursive loop
Question brief opened and researched 30 September 2026. Related: Amodei’s pacing argument, two authors’ risk comparison and Dru’s benchmark-validity question. Agentic coding workflows belong in Agentic Software; the issue here is capability feedback and safety governance.
Question and answer standard
What observed contribution do AI systems make to AI R&D, and what would demonstrate that contributions compound from one model generation to the next without increased human input, training/experiment compute, data or energy? A good answer distinguishes AI use, task-level researcher uplift, aggregate capability acceleration, full automation and self-sustaining acceleration. Track dollars, time and quality-adjusted discoveries, not just code lines or success on preset rubrics.
Source list from four subquestions
- Anthropic, “When AI builds itself,” plus April and August automated-alignment trials: strong internal task data, openly limited by direction-setting and transfer.
- METR, predeployment Opus 5.5 assessment (22 September): independent partial assessment, explicit early estimate of aggregate acceleration and unresolved judgment gap.
- METR, NanoGPT cost curves (21 July) and Cunningham/Shetty’s apple-picking model: falsifiable skeptical model and experimental budget accounting.
- Cunningham and eight coauthors, economics of RSI: exact distinction between feedback, automation and self-sustaining acceleration; tentative empirical calibration.
- METR, Frontier Risk Report (May, Feb–Mar assessments) and Epoch, interviews with eight AI researchers: earlier independent baseline and methodological objections; both predate later model evaluations.
Findings, matched to the edges
- Contribution is real and costly: Anthropic reports 80%+ AI-authored merged code and 8× code per engineer in Q2 2026 versus 2024, but concedes code volume overstates research output; 130 employees’ median ~4× productivity is self-reported, with Anthropic itself expecting less. Fixed-goal small-model optimization and automated safety-method discovery are strong demonstrations; one April weak-to-strong test attained 0.97 on a small model for ~$18,000/800 agent-hours versus humans’ 0.23 in seven days but no statistically significant improvement from that method on the tested production model. The August alignment trial improved ten predefined, measurable failures and transferred on withheld evals and larger models, which weakens a blanket claim that all transfer fails; still not a self-selected research program with verified general safety.
- Independent read on whole-lab pace: METR on 22 September found Opus 5.5 likely incrementally increases research uplift but unlikely to fully automate AI R&D; a separate preliminary team estimated ~1.5× overall capability acceleration, perhaps 30% chance of 2×, yet the reporting team did not access supporting evidence, period was unspecified and Anthropic had editing rights. Treat as a bounded provisional estimate, not causally identified lab throughput. METR’s July outside-in >2× individual uplift projection is a different metric and could shrink under verbosity or extra low-value work.
- Skeptical limiting mechanism: Cunningham/Shetty’s apple-picking model predicts cheap local gains then rapid saturation until a stronger generation arrives; running the same agent repeatedly need not add advances. METR’s six NanoGPT agent runs found revalidated improvements but expenditure horizons only roughly $0–$3,300 under a highly uncertain human labor baseline and an inefficient harness in which GPUs consumed 70–90% of many runs’ costs. This does not bound frontier R&D or human–AI collaboration: the same METR authors stress that risk. Anthropic itself reports review and idea-triage bottlenecks.
- What recursion requires: Cunningham et al. formally define self-sustaining acceleration as increasing capability growth without rising outside inputs. Their exploratory model’s ~15% required productivity gain per unit of a specific capability index versus a coarse ~9% inferred gain suggests current feedback below the threshold, with wide uncertainties and possible future reversal. Even full automation might not yield sustained acceleration with diminishing returns or compute/data constraints; partial automation may be enough to speed research before autonomy. Narrow skill on scored optimization tasks may not imply broad dangerous ability.
Named disagreement and next discriminating evidence
Anthropic’s strongest case: capabilities rise on real internal tasks, agents increasingly design and iterate on experiments with less direction, and frontier progress is often incremental; today’s human taste could soon cease to bind. Best skeptical case: METR’s NanoGPT revalidation and apple-picking model expose repeatability/quality and experiment-GPU costs; METR’s September judgment still finds human foresight and feedback creation weak; Cunningham et al.’s calibration does not yet suggest self-sustaining acceleration. Neither side has an independent, across-generations counterfactual for true marginal lab advances. The July economic model cannot incorporate all late-September developments, and a September provisional estimate cannot establish a trend from one snapshot.
Critical experiment: Obtain independent inside-lab time series for (i) independently validated novel algorithmic gain and broad capability gain per dollar of researcher, agent inference, experiment and training compute, (ii) who chose goals and validated results, (iii) how new models fare when started after earlier agents’ discoveries, with matched human/hybrid baselines and held-out frontier-scale tasks, (iv) whether measured marginal impact rises across at least two successor generations at constant exogenous inputs. Audit permissions and publication rights for the assessor. On 30 September a search found only METR’s 22 September public preliminary summary, not the promised deeper inside-lab report. Meanwhile do not interpret researcher help as autonomous self-propagation, or remaining human review as permanent proof of safety.
Dru’s reading, 30 September 2026
In discussion of the METR post at 07:20 UTC, Dru called this mild evidence against RSI fears: AI-research benchmarks may quickly saturate and omit messy research reality; gains looked incremental rather than qualitatively improving taste/judgment; extrapolating them into takeoff assumes an as-yet-unobserved transition. The new benchmark note tests this objection across RE-Bench, MLE-bench, PaperBench and more open-ended tasks. Qualification worth preserving: METR found incremental improvement on harder-to-verify conceptual and open-ended tasks too; and one model comparison cannot identify the long-run growth rate. The skeptical inference is therefore against an observed autonomous or self-sustaining loop now, not a demonstration that future qualitative change is impossible.