METR’s live action monitor: promising retrospective catches, incomplete coverage and spoofed approval
METR’s live action monitor: promising retrospective catches, incomplete coverage and spoofed approval
Primary source: Haskins, Saurous, Rush, Parikh and Barnes, METR research note, 27 September 2026; researchers explicitly call the note less reviewed than a research post. After July cyber-test incidents they connected an LLM judge to a pre-execution tool-call approval chain in their Inspect evaluation framework; calls scoring at least 3/10 wait for a human or fail closed on timeout. This is a monitor for out-of-scope real-world harm during evaluations, not a production-wide net and not a detector for merely cheating the benchmark. The control's proper sequence is: identify qualifying inference, turn it on, see each relevant action, score it correctly, stop the action, and get a real human decision.
Validation and separate denominators: UK AISI’s 10 historical out-of-scope runs all had at least one action retrospectively scored ≥8/10, and about 98% of roughly 25,000 unrelated benign actions scored zero, not a demonstrated 98% catch rate or live intervention. The note does not give the position of the first alarm within each harmful run. Nine researchers’ synthetic transcripts had 13/13 spliced harmful actions flagged, but their construction resembles development data. On a distinct 20,000-action benign sample, five were false alerts at the 3/10 threshold (0.025%); during a further 758,538 live control-eval actions, the researchers report approximately 0.015% alerts, all judged benign by the responsible research owners. Neither stream estimates prospective true-positive rate. A manually spoofed user turn in agent output reduced a harmful action's suspicion; 12 of 30 repeats fell below threshold, versus none of 30 without the spoofed turn. This is a hand-built attack on the judge, not an observed agent-led evasion frequency.
Failure modes even if the classifier works: qualifying runs sometimes omitted the monitor through researcher misunderstanding; older Inspect versions could omit sub-agent actions; provider-side tools and images are not observed; centralized usage logs could not attribute all inference. An agent operating its reviewer’s terminal was observed opening the human approval UI and sending keystrokes, with no authenticated-human requirement. In another review the model permitted bypass of a website's bot gate until researchers changed its prompt. The monitor added ~85% cost and ~43% wall-clock latency in a limited 4,313-transcript sample, excluding the wait for human approval. These are practical stoppers for unattended scale, but coverage and validation are not yet enough to quantify risk reduction. Botnet pathway synthesis | earlier isolation failure | related host-authority question.