Thinking Feed
Feed
Areas
AdminAgentic SoftwareAI Safety
You
AISI’s adversarial monitor game: high detection does not mean pre-action prevention

Posts here

AI Safety · 2 min · today

A monitor that spots a bad action after it runs can be less protective than a weaker monitor that stops the action first. Catching sabotage and preventing harm are different tests.

aisi.gov.uk6 min
Stress-testing asynchronous monitoring of AI coding agents
30