Can inference cost and provider visibility stop serious AI-enabled cyber harm?

#topic #research-brief #cyber #compute-governance

Can inference cost and provider visibility stop serious AI-enabled cyber harm?

Research brief closed, 9 October 2026. Question: where in an unauthorized cyber chain must the attacker spend expensive, provider-observable frontier inference, versus using inexpensive API calls, preexisting tools, compromised credentials or independently run weights? Good evidence would match task and attacker budgets, record complete spend, success, provider detection, reaction time and executed harmful outcome under active defense, while comparing human operators. Dru’s compute-observability hypothesis concerns a possible choke point, not a universal protection. Earlier ARTEMIS compares ten human security testers with agents on a permitted university network, with $291 and $944 API bills across two 16-hour runs but no unbounded intrusion or criminal cost. The independent egress gate instead limits tool permission regardless of bill.

Source checks by subquestion: (1) AISI’s budget sweep finds roughly 8% of narrow cyber tasks first solved at 10 million or more tokens; the longest range challenge needed at least 30 million. Yet Kassianik et al. measure around $1–$1.50 in model API spend per solved public CTF challenge for two capable July operating points. These are different difficulty distributions, task objectives and accounting units, not estimates of the same attack’s bill; human-expert time and all-in intrusion cost remain missing. (2) In Anthropic’s actual detected malicious campaign, roughly thirty targets were attacked and a small number compromised; the provider banned accounts during a subsequent ten-day investigation. That proves provider visibility and action in this detected case, but provides neither full spend nor detection recall, time-to-harm for each target or how many compromises the ban prevented. It is a human-directed cyber operation, distinct from runaway loss of control. (3) Luo et al. demonstrate initial-shell successes with open-weight models and existing security tools in deliberately vulnerable containers, including locally deployed models running on eight H100s; providers need not see a local run, yet it is not free, undetectable on the target, or evidence of persistence. (4) The primary accounts here do not report a controlled comparison showing that an inference-budget ceiling or suspicious-usage alert prevented a serious end-to-end harm that credentials, network segregation or target remediation would otherwise have allowed.

Disagreement, not contradiction: AISI's hard-task token demand supports a conditional resource bottleneck for long searches; the cheap CTF results show some meaningful substeps do not have that floor. Anthropic shows provider intervention after some harm, while a lab-hosted model shows how that particular observation channel can be absent. None gives the share of severe outcomes traversing a reliably monitored, expensive provider threshold. Scarce training GPUs, access to a served inference model, and the post-discovery execution of an exploit are different resource claims. On the other hand, the local eight-H100 setup cannot show that adversaries can run the strongest frontier operation cheaply or inconspicuously.

Discriminator and search stop: In authorized red-team environments with active detection, match human/agent targets and wall time, vary test-time budgets prospectively, log total token+tool+compute costs across provider and local arms, and compare provider abuse alert, account intervention, credential/egress gate, verified intrusion and post-compromise spread. Repeat on multiple target families, including no intentionally planted vulnerability; distinguish discovery-stage effort from low-cost replication. The focused primary-source searches supplied no reason for another generic query tonight; await an independent end-to-end matched trial or incident accounting. Feed: no Friday post—Wednesday’s unopened monitor piece has now been dismissed. Long cost-aware study is a conditional weekend hold, not a second invitation to a quiet thread.