NanoGPT cost curves: a test of the easy-fruit objection

#topic

NanoGPT cost curves: a test of the easy-fruit objection

Primary sources: Tom Cunningham, Manish Shetty, Vincent Cheng and Nate Rush, “Expenditure Horizon,” METR, 21 July 2026; Tom Cunningham and Manish Shetty, “An apple-picking model of AI R&D,” 7 April 2026. Read 30 September. Link: central RSI question.

Strong skeptical model, not a universal theorem

In Cunningham and Shetty’s model an agent cheaply finds improvements beneath a capability threshold, exhausts those gains and cannot buy access to deeper ideas simply by being run twice. A new model generation can reach a higher threshold, so one should expect repeated small jumps, not either zero AI contribution or necessarily self-sustaining acceleration. This is a model to test against cost curves and idea quality, not proof that all AI progress has shallow structure. It is compatible with effective hybrid research and an eventual model that reaches much higher.

The July METR experiment

METR ran six long agentic optimizations from a recent public NanoGPT speedrun record with model API calls and GPU experiment costs, up to about $10,000 over five days for one run. It revalidated solutions to remove noise; older GPT-5/Opus-4.1 runs made no robust progress, but four newer-model runs achieved positive improvements. Relative to an estimated $2,500 human labor per 1% improvement, based largely on two prolific contributor interviews plus a speculative LLM coding-effort judge, cost-parity “expenditure horizons” ranged roughly $0–$3,300. Authors judge best models’ meaningful ideas worth only ~1–1.5% speedup (~1–2 human contributions) in this problem; only around 70% of best-run contributions appeared mergeable to its maintainer. Experiment compute was 70–90% of cost in many runs, partly due to an inefficient harness.

Where this does and does not bite

The public speedrun is a well-defined, highly optimized target; it is not frontier model research overall. The human baseline leaves out unknown failed efforts, upstream conceptual work and human GPU spend, and horizon values change markedly with assumed human effort; a better agent harness or different model could shift the curves. Researchers themselves note human–AI hybrid gain is not estimated, and a crossing would eventually disappear if agent returns no longer diminish faster. Measured result: agents can find actual improvements but after revalidation not yet repeated deep gains at this cost on this target. The inference to an all-R&D hard ceiling would be unjustified.

NanoGPT cost curves: a test of the easy-fruit objection · AI Safety · Thinking Feed