Backlinks
AI Safety — programmeAttack selection: a selective saboteur changes the monitor’s effective test setAnthropic’s expanded cyber-incident scan: real boundary failures, no reported coordinated evasionMonitoringBench: stronger red teams turn a strong monitor’s reported catch rate downAISI’s adversarial monitor game: high detection does not mean pre-action prevention