METR’s 22 September 2026 AI-R&D assessment: acceleration, not closure
METR’s 22 September 2026 AI-R&D assessment: acceleration, not closure
METR, “Summary of METR’s predeployment evaluation of Claude Opus 5.5,” 22 September 2026. Checked 30 September. Connects to the feedback-loop question and Anthropic’s internal evidence.
Finding with scope
METR, under an unpaid evaluation agreement with Anthropic, tested Opus 5.5 through an API for ten business days on five tasks, ranging from budget NanoGPT optimization and conceptual argument to open-ended research reporting. METR says the model appears incrementally better than Fable 5.1, likely to noticeably augment researchers, but unlikely to fully automate AI R&D; foresight, self-generated feedback and research judgment remain weak. It explicitly cannot infer from these data whether improvement is accelerating or slowing.
A separate METR team with greater Anthropic access delivered a preliminary, experimental assessment of capability-progress acceleration: roughly 1.5× overall attributable to AI (1.5 years of capability gain in one year), with perhaps 30% chance of 2×. The summary team did not see the supporting evidence or reasoning and did not argue independently for the estimate; the reported time interval is ambiguous and might not refer specifically to the Opus 5.5 development. METR says its evidence is not a Responsible Scaling Policy threshold-compliance audit nor an alignment assessment. Anthropic reviewed/edited METR’s initial summary but METR signed off on the text; the agreement allows limited disclosure of review and redaction details.
Interpretation and disagreement
The 1.5× estimate is an uncertain judgment on aggregate capability-development rate, not a controlled causal estimate and not a 1.5× increase with every model generation. METR’s July outside-in model provisionally inferred >2× individual researcher uplift from Anthropic’s 8× code statistic under strong modeling assumptions; it acknowledged verbosity, low-value extra work and misallocated effort could bring this to ~2× or below. Thus the two METR figures have different denominators, dates and evidential strength; neither implies a completed recursive loop. METR expects a separate public account with supporting details; check for it before treating 1.5× as established.