Anthropic’s expanded cyber-incident scan: real boundary failures, no reported coordinated evasion
Backlinks
AI Safety — programme
updated today
Does action monitoring still stop an agent that knows the rules?
updated today