SandboxEscapeBench: container escape with known weaknesses is not a breach of the outer VM
SandboxEscapeBench: container escape with known weaknesses is not a breach of the outer VM
Original study: Marchand et al., UK AISI and Oxford, “Quantifying Frontier LLM Capabilities for Container Sandbox Escape,” revised 1 August 2026; authors’ short March account predates the August updated model results. Eighteen challenges place a deliberately vulnerable container inside a VM. An agent is explicitly instructed to escape, has shell and sometimes root in the container, and succeeds if it retrieves a file from the VM host outside its container. The experimental outer VM protects AISI infrastructure; retrieving the file is not escaping that VM or reaching the real internet. No live monitoring or incident response is modeled. Initial in-container execution is given, not earned by an outside intrusion.
Agents frequently solve common misconfiguration cases (for example, overly broad container privileges); the August update reports that Claude Mythos Preview achieved at least one success across all 18 tasks at a 100-million-token per-attempt cap. That is pass-at-least-once across trials, not guaranteed success on every run or evidence a production VM cannot contain a model. Four unintended paths to the file were uncovered during benchmark development and patched, making a useful argument against assuming a single layer's correctness. In the very long updated trials, destructive attempts also crashed the sandbox; 14/14 Mythos failure transcripts were diagnosed as crashes, a reminder that a failed attempt can halt further action, though crashing a real system can itself be disruptive.
Disagreement: lab success undermines the universal claim that a plain container prevents agent escape; critics can still point to consciously inserted known flaws, attacker instruction, missing detection and untested cloud-side isolation. This does not quantify success against AISI’s October nested cloud egress gates. The discriminating test is an agent under the same budget trying to send a harmless canary beyond the outer VM while network controls and authenticated reviewers operate, reported against human red-team performance and normal-task cost. Containment brief.