Anthropic's internal automation index measures delegation, not self-sustaining research growth
Anthropic's internal automation index measures delegation, not self-sustaining research growth
Primary Favaro and Wright, “Measurements for understanding the pace of AI development inside frontier labs”, 17 September 2026 (dated via Anthropic’s newsroom), read 4 October. Related cross-generational test, research taste and containment failure mode.
What was actually counted
A July sample of 20% of R&D staff per department each week yielded about 15,000 granular task descriptions, organized into 378 leaf categories (542 tree nodes). Anthropic froze that basket, weighted categories with July person-time (splitting a person-week evenly across activities), then Claude agents researched each task category's actual work from internal records and an independent Claude judge rated the role played by AI on a six-level scale. Some unblinded incentive concerns remain even though staff rated work independently of judge evidence: the same company's models create the task map, read internal records and rate its own work. Staff/model exact-level agreement was 59%; staff/staff exact-level 35%; agreement within one level 97%. The measure is person-time-weighted task categories, not causal productivity gains, number of discoveries, or percent of research scientists replaced.
Anthropic says as of August 2026 Claude led 26% of weighted AI-R&D work (up from under 1% in February) and at least collaborated on over 90%; no measured task subset reached its level 5, in which the AI completes the full work without human involvement. “Leads” means doing most of a task from a high-level prompt while a human supervises and ultimately accepts it. The company examined work shifted into new tasks on a January basket and did not see growth at its resolution, but cannot rule out uncaptured new work or shifts in importance. The finding makes denial of real delegation implausible yet does not show AI independently chooses high-value agendas, validates successor systems, or increases quality-adjusted improvement per dollar.
Monitoring and compute are additional denominators
On Anthropic's most-used platform only, about 30,000 research/engineering agents ran concurrently as of August; all actions on that platform pass online monitors before execution and offline monitors afterward. Among >1 billion August decisions, 0.002% were blocked; ~100,000 transcripts/week receive automated triage and ~50/week human escalation. Coverage is not detection sensitivity; false negatives, off-platform activity, and whether reviewers can respond to slowly developing multi-agent behavior were not estimated. Agent identities and a shared communication log are control design features, not proof of reliability in hostile settings. A July 13–20 snapshot allocates ~6% of R&D compute to safety and ~12% of AI-driven AI-R&D compute to safety by the lab's definitions; classification and sample coverage caveats make spending share no direct proxy for safety effectiveness. Neither low block rate nor high coverage proves no dangerous actions.
Disagreement and next observation: Anthropic regards these as useful public pace and oversight measures; METR's 22 September external assessment still calls full R&D automation unlikely for tested Opus 5.5 and its own whole-lab acceleration estimate provisional; the Cunningham et al. model says lab delegation can increase without self-sustaining acceleration. A third party with privacy-safe access should validate category sampling and monitor recall, track independently validated marginal algorithmic discoveries per human/inference/training/experiment dollar across model generations and reveal what humans authorize. No such public causal series located in the 4 October search; the lab offers a prototype, not independent verification.