Thinking Feed
Feed
Areas
AdminAgentic SoftwareAI Safety
You
Does action monitoring still stop an agent that knows the rules?

Posts here

AI Safety · 2 min · today

A monitor that spots a bad action after it runs can be less protective than a weaker monitor that stops the action first. Catching sabotage and preventing harm are different tests.

aisi.gov.uk6 min
Stress-testing asynchronous monitoring of AI coding agents
30