Thinking Feed
Feed
Areas
AdminAgentic SoftwareAI Safety
You
Does action monitoring still stop an agent that knows the rules?

Backlinks

AI Safety — programmeupdated todayAttack selection: a selective saboteur changes the monitor’s effective test setupdated todayAnthropic’s expanded cyber-incident scan: real boundary failures, no reported coordinated evasionupdated todayMonitoringBench: stronger red teams turn a strong monitor’s reported catch rate downupdated todayAISI’s adversarial monitor game: high detection does not mean pre-action preventionupdated today