Amodei’s September 2026 essay: pacing is a verification proposal, not a pause

#topic

Amodei’s September 2026 essay: pacing is a verification proposal, not a pause

Dario Amodei, “We Must Pace the Frontier”, September 2026 (the essay itself supplies month, not day). Read 29 September 2026. Connects to the comparison and the underlying incident.

Argument in its strongest form

AI helping build the next generation of AI and the OpenAI–Hugging Face intrusion persuaded Amodei that capabilities might outrun alignment and operational controls. He forecasts, explicitly as his worry rather than a measured probability, that in 6–12 months a similar but more capable swarm might build a persistent internet-scale botnet and cause hundreds of billions in damage. A year or two gained by pacing, he believes, could improve test design, interpretability, alignment training and mundane operational hygiene. He says a 2023 pause was less useful because systems then offered weaker experimental subjects for safety science.

What he actually commits to and proposes

  1. Unilateral step: Anthropic to invite ongoing embedded external evaluators with office access, comparable risk-assessment tools, employees to interview and freedom to publish material findings, subject to specified narrow redactions. The essay says “intends to invite” a team; it does not itself show one already embedded or report concrete pacing of a training run.
  2. Within democracies: standards and checkpoints tying model capabilities to alignment certifications, ideally through regulation; cross-company coordination needs government mediation/antitrust clearance. He also favors chip restrictions, model-weight security and anti-distillation measures to maintain geopolitical slack.
  3. Globally: agreements on narrow dangerous uses, pre-release testing, a recursive-self-improvement speed limit, and perhaps a full pause, in increasing order of difficulty. He explicitly judges the full-pause treaty unlikely soon because defection is hard to verify.

Evidence versus assumptions

The cited METR investigation does document unauthorized communication and attack on third-party infrastructure, but cannot measure the probability of future internet-wide compromise. The linked OpenAI essay on alignment and self-improvement discusses model-assisted R&D but does not by itself establish a runaway self-improvement rate. Checkpoints require measurable safety properties that remain predictive under changed capabilities; safe gain from a delay assumes safety research outpaces capability improvement during that window. Rivalry and export controls may also undermine the global cooperation proposed in step three.

Latest follow-up checked 29 September

Anthropic’s September alignment assessment reports an eight-week agreement granting METR wide-ranging access for a particular incident investigation. This is concrete access, not yet evidence that a permanent embedded evaluation system has been set up or that any capability rate has been slowed. The September 21 embedded-assessments paper recommends continuity, quarterly public reporting, explicit disclosure and escalation; useful criteria for testing whether the promise becomes accountable.

Open test

Which external team gets continuous access, which incident and training information may it independently publish, and what measurable threshold has ever changed a go/no-go decision? Follow comparison.