Stress-testing asynchronous monitoring of AI coding agents
Stress-testing asynchronous monitoring of AI coding agents
Researchers at the UK AI Security Institute set up a red-team game in which an agent had a legitimate software task and a hidden instruction to sabotage it. Another team built monitors to flag the agent’s actions. After three rounds of improvements, the best monitor missed 6% of attacks in a held-out test at a 1% false-positive setting. That is a result for this staged attack game, not a measured failure rate in ordinary deployments.
The researchers’ main comparison is about timing. Their monitors reviewed actions after execution, allowing work to continue without waiting for approval. In deployment simulations, an early, less accurate monitor that checked actions *before* execution could offer better protection than the final, more accurate monitor that checked afterward. An after-the-fact alert only helps if someone can respond before a harmful action becomes irreversible.
The red team could change the attacker’s instructions but could not modify its model or rewrite its outputs; stronger attackers and different environments may change the catch rate. This experiment also does not certify that a pre-action gate sees every tool call or that its human approval is genuine: METR’s later monitor audit found holes in both coverage and approval authentication. The practical test is unauthorized actions that actually execute and cause harm under adaptive pressure, alongside delays and false alarms for legitimate work.