Economics of recursive improvement: a threshold, not a prophecy
Economics of recursive improvement: a threshold, not a prophecy
Primary: Tom Cunningham et al., “The Economics of Recursive Self-Improvement,” dated 13 July 2026 (arXiv submitted 14 September 2026); authors’ METR explanation, 22 July. Read 30 September. Linked synthesis: RSI evidence versus bottlenecks.
Precise question
The authors avoid using “RSI” as a yes/no because any assistance to AI R&D can be called feedback while full autonomy and self-sustaining acceleration are distinct. Their operational target is acceleration of capabilities without growth in exogenous inputs, such as human researchers or training compute. They explicitly argue full automation need not imply such acceleration if new discoveries get harder or complementary inputs remain scarce. Conversely partial automation can strengthen the feedback before humans leave the loop.
Model and calibration
Their directed graph connects AI capability → effective research effort / algorithmic improvements → future AI capability; loop strength depends on products of elasticities, with training and experiment compute, inference, data and human judgment as potential bottlenecks. A narrowly better optimizer need not produce broadly more capable or more economically influential AI. Their tentative back-of-envelope calibration uses the Epoch Capabilities Index: an index-unit increase would need to lift AI R&D productivity by at least ~15% under their assumptions, versus an inferred ~9% from reported engineer uplift since coding agents. These are index-normalized model values with wide uncertainty, not direct controlled estimates or safety thresholds. They conclude existing feedback likely below self-sustaining level, possibly strengthening; they do not rule out crossing the threshold soon.
How to test and disagreement
The most uncertain edge, say the authors, is capability’s effect on algorithmic discovery rate. They ask labs for measured algorithmic efficiency, R&D spending split across researcher labor, inference and experiment compute/data, and traceable AI contribution to actual technical advances. Contrast Anthropic’s evidence of fast experiment iteration with METR’s cost-curves on a limited optimization problem and METR’s September uncertain acceleration judgment. A truly discriminating test would track successive generations’ marginal independently validated discoveries, quality-adjusted capabilities and compute used with exogenous inputs separated; repeated small success on known rubrics alone does not identify the feedback elasticity.