Provider disruption after an AI-assisted espionage campaign is not a prevention-rate estimate
Provider disruption after an AI-assisted espionage campaign is not a prevention-rate estimate
First-party incident account: Anthropic, 13 November 2025, “Disrupting an AI-orchestrated cyber espionage campaign”. Anthropic says its abuse team noticed suspicious Claude Code use in mid-September 2025, later attributed it with high confidence to a Chinese state-linked actor, and mapped attacks on roughly thirty targets; a small number were compromised. It reports banning accounts as identified during a ten-day investigation and notifying affected parties. The human adversary chose targets, built and steered a framework, and intervened at several decision points; the provider estimates the model handled 80–90% of the work. Those proportions are its reconstruction, not an independent time-and-motion comparison with human hackers.
This is direct evidence that a cooperative provider can observe suspicious malicious use and take action, and that provider controls did not stop all reported compromises before discovery. Anthropic does not supply complete API credit cost, threshold to identify the operation, time-to-detection for each target, false-positive/false-negative rate, or the counterfactual harm avoided by the bans. The disclosed actor used a commercial model service and accounts; observation of one detected actor is not a denominator of all users, especially self-hosted models. This is human misuse of an agent, not Anthropic’s four 2026 third-party eval agents independently reaching real services and not an internet-spreading rogue botnet. See cost/visibility test and cyber outcome chain.