Measurements for understanding the pace of AI development inside frontier labs

#summary

Measurements for understanding the pace of AI development inside frontier labs

Anthropic proposes publishing three kinds of internal evidence about how AI gets built: which research tasks AI performs, whether agents’ actions are monitored, and where research compute goes. Its prototype rates task categories using company records and model judges, weighted by staff time. By August, it classified 26% of its AI research and development work as AI-led, up from under 1% in February, while finding no measured category that AI completed without a human in the loop. “Led” means the model does most of a task after a high-level instruction, with a person supervising—not that it selects the lab’s goals or produces a validated successor independently.

The same piece reports about 30,000 agents running concurrently on its main internal research and engineering platform. Every action on that platform goes through an online monitor, but the fraction blocked is not a measure of how many dangerous actions the monitor misses. A one-week compute audit classified about 6% of AI R&D compute as safety work; spending share cannot measure whether that work succeeds.

Anthropic argues that regularly published, independently checked versions of these measures could make rapid development and oversight legible to outsiders. For now this is company-run measurement: Claude helped create and judge the task map, the agent monitoring figures cover only one platform, and neither the task share nor the compute snapshot establishes quality-adjusted research acceleration across successive models. An independent assessor would need validated advances, human review and all-in experiment costs across generations to test that stronger claim.

Read at anthropic.com · 14 min