Can measured research assistance compound into self-sustaining AI improvement?
Can measured research assistance compound into self-sustaining AI improvement?
RSI synthesis | Benchmark-validity objection | Economic threshold model. Weekend research brief completed 4 October 2026. Question: is there evidence that agents choose promising problems, validate expensive advances and raise their next generation’s marginal research productivity? A good answer separates preset objective iteration, research triage, actual lab delegation and independently validated quality-adjusted across-generation progress; counts human decisions, training and experiment compute, and does not equate faster research with loss of control.
Queries and source shortlist
- Original lab experiments on autonomy and transfer: Anthropic’s August automated alignment research; TASTE original paper.
- Feedback thresholds: Cunningham et al., economics of RSI, and coauthors’ short note.
- Independent scoreability/cost: METR's September model evaluation and July expenditure horizon.
- Observed labs rather than testbeds: Anthropic’s September internal automation, oversight and compute prototype; Chan et al.’s measurement proposal (proposal, not collected dataset). The search did not locate an independent cross-generation causal quality-adjusted lab discovery and cost series.
What each edge now says
Local exploration → verified proxy improvement: Anthropic's five-agent, GPU-capped alignment experiment researches literature and hill-climbs ten specified benchmarked safety failures. Top methods improved held-out benchmarks and models up to 4.7× larger; 30 human-proposal-seeded runs were not better than 30 without human seeds. This is the strongest rebuttal to “agents cannot generate useful research ideas without a human-proposed method.” But humans specify the failure, design evaluation gates and authorize deployment; human baseline got one noniterative proposal while agents tried many; 39 of 1,601 method traces triggered a cheating audit. The authors concede no test of rare unknown failures, large posttraining persistence or general research agendas. Original full report.
Open-ended triage → good research choice: TASTE found its best frontier model predicts selected experienced-human proposal preferences on 60% of 92 comparisons, versus estimated human–human 77%; expert labels themselves have substantial noise, and preferences are not downstream project success. The model improves with inference effort and performs better on cross-topic pairs. This supports Dru’s current research-judgment gap, not a permanent ceiling; it cannot refute success on a different narrower feedback loop.
Deployment → lab pace: Anthropic estimates AI led 26% of person-time-weighted task types by August (less than 1% in February), at least collaborated on over 90%, and independently completed none of its measured types without human involvement. This is an in-house model-judged proxy for delegation, not a measured marginal impact on breakthroughs. METR separately estimates some lab acceleration but not full R&D automation on its tested model, with no disclosed supporting data for its preliminary whole-lab estimate. The lab result and external assessment are compatible rather than contradictory: supervising most tasks still limits full autonomy.
Feedback → self-sustaining acceleration: Cunningham et al. require sustained increases in quality-adjusted capability progress at fixed external inputs; preliminary ~15% per-index-unit threshold and ~9% guessed productivity response are uncertain, differently sourced model-dependent numbers, not a safety pass/fail. The NanoGPT cost curves show limited repeated gains and meaningful experiment spend on one public optimization target. Lab-scale success could outrun that problem; conversely a growing task-automation share could fail to raise capability growth when humans, data, GPUs, energy or slow evaluation bind.
Named disagreement, next test, feed
Anthropic’s strongest evidence for fast feedback is real task leadership and held-out transfer after many machine-run iterations, including unguided method search. METR’s best outside read finds only modest per-model gains and no observed large increase in independent foresight; TASTE shows a present expert-agreement gap and Cunningham et al. show why task delegation is not logically enough for autonomous growth. Missing observation: an independent assessor accesses successive real lab projects, pre-registers hard-to-verify project outcomes and marginal verified capability gains, and reports human agenda/review time, inference/GPU experiment/training costs, new tasks and failure/monitor recall across at least two successor generations. Similar or growing improvement per dollar at fixed outside inputs would change the skeptical case; mere rising delegation would not.
Feed choice, Sunday: Anthropic's authored measurement essay for a single accessible but substantive index post; TASTE’s 61-page original paper held for a later weekend if the index or judgment gap is discussed/hearted. The economics paper remains a later long-read if Dru asks to derive the threshold; do not pile three unopened RSI studies on an already quiet feed.