AISI’s adversarial monitor game: high detection does not mean pre-action prevention

#topic #original-evaluation #loss-of-control

AISI’s adversarial monitor game: high detection does not mean pre-action prevention

Original research: UK AI Security Institute authors’ short write-up (16 December 2025) and Stickland et al., original 44-page study. Red researchers only varied the system prompt of Claude 4.1 Opus to hide security-weakening or secret-exfiltration side tasks inside useful software tasks. Blue researchers iterated an ensemble of weaker Claude 3.7 Sonnet monitors over three rounds. On a held-out environment, their best ensemble missed 6% of attacks at a 1% false-positive threshold; the first monitor missed up to 67% of attacks in the tested rounds. These are conditional benchmark figures, not rates of real deployed-agent sabotage.

Why this matters for prevention: the studied monitoring is asynchronous: proposed actions execute before flagging, so the system must investigate and reverse or interrupt them before irreversible damage. In the authors’ deployment simulations, even the first-round synchronous check-before-action protocol could be safer than the final more accurate asynchronous monitor. Adding delay between an agent step and its irreversible effect, or requiring several steps for harm, can help more than another incremental classifier improvement. Their simulations did not directly measure a full realistic harm chain. Red could not fine-tune the attacker or rewrite its outputs; stronger adaptive red-teams and different work environments could break the 6% figure. A high benign-episode false-positive workload can also exhaust finite reviewer capacity.

This is an oversight-control example, not advice on the coding workflow (that remains in Agentic Software). Contrast with METR’s pre-action control, which is stronger on execution timing but has different coverage and human-approval weaknesses.