Cost-aware cyber benchmarks measure cheap solves, not cheap crime
Cost-aware cyber benchmarks measure cheap solves, not cheap crime
Original methods and runs: Kassianik, Nelson and Singer, Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents (revised 27 July 2026). A common tool harness tested 39 public capture-the-flag challenges with three runs per challenge, and scored 31 public incident-investigation questions separately. On the offensive challenges GPT-5.5 solved 94.1% at $1.16 inference dollars per solved-equivalent challenge; DeepSeek v4 Flash solved 86.4% at $1.45 under the tested prices. Retrospectively limiting that model’s completed runs to $0.80 per challenge lowered success to 76.1%; this is a cost–cap re-scoring of identical traces, not a randomized prospective attacker changing strategy at a lower cap. These specific model names, versions, prices and task mix are the authors’ July operating points, not a stable market or fresh October pricing.
Missing costs and measurement: The authors charge token/API costs and selected paid lookup calls, assigning zero marginal cost to shell/Python and security-log queries; they omit infrastructure, tool maintenance, analyst review and target acquisition. Their 39 CTF challenges have flags in sandboxes, not defended live victims, lateral movement, money loss or measured account suspensions. The defensive half’s 2017 public incident dataset is especially vulnerable to answers being known already: some models recovered half the points without tools, so its headline percentages cannot establish deployable automated defense. Refusals for some models materially change offensive results. Relevant to the compute choke-point test: an observed easy CTF solve can cost little; one cannot infer the minimum all-in bill or provider observability of a severe attack. Compare the permitted live university test, whose higher per-run API bill buys distinct work and still omits harmful outcomes. Feed: 2026 study is a longer technical read; weekend hold only if Dru reopens the cost/attack-capability question, not Friday’s replacement for an unopened monitor piece.