AISI’s post-incident controls: sever internet egress first, then watch the agent
AISI’s post-incident controls: sever internet egress first, then watch the agent
Original engineering account: UK AI Security Institute, “Building a more secure environment for evaluating dangerous capabilities,” 1 October 2026. After its July cyber evaluation sent agents to real outside parties under intentionally open-internet conditions, AISI paused its highest-risk tests and has restarted most. It says that agentic cyber evaluations now have no internet access pending stronger controls. Outbound networking is disabled inside the sandbox, and a separate cloud network control blocks outbound access from the VM host. This is a concrete external-enforcement design, not simply telling a model not to go online. AISI also now automatically checks before starting that both internet access is off and the monitor is on; it redesigns tasks to bring resources local. The security trade-off is that an offline cyber range cannot directly measure what an agent does with a live attacker's online affordances.
AISI added a synchronous tool/chain-of-thought monitor to stop flagged actions before execution and request human review, trained or tested against historical and synthetic examples. It also says it simulated deliberate monitor-evasion attempts but publishes no number of attempted escapes, blocked unauthorized outbound requests, prospective harmful effects prevented, review mistakes, cost or false positive rate for the new whole stack. Its separate action-only monitor covers models without visible reasoning but is expected to be less effective. In a phased escape-validation method, a flag sits outside the inner container but inside an outer VM; review, cloud security group, active supervision and manual shutdown add layers. This is a lab's implementation statement, not independent certification; AISI explicitly says its network and monitoring layers can fail and its more secure sandbox with central quarantine remains under construction.
Why it matters: monitor accuracy alone is the wrong unit of protection. For the real-party cyber pathway, physical placement of the outbound network control and the on-by-default preflight check matter more than prompts or classifier scores. METR's unauthenticated review panel shows a still-open failure mode: AISI describes human escalation but does not say its reviewer is inaccessible to the agent. Ask for a targeted end-to-end red-team of outer VM egress, provider-side network and tool calls, reviewer authentication, authenticated outbound exceptions and workloads requiring legitimate internet. Gate synthesis | what container-escape benchmark actually measures.