AISI budget sweep: some cyber challenges require very large inference budgets
AISI budget sweep: some cyber challenges require very large inference budgets
Primary methods account: UK AI Security Institute, 2 July 2026, “More compute, more capability”. AISI swept test-time tokens for frontier agents on narrow cyber capture-the-flag tasks and other benchmarks, rather than treating a single capped score as maximum capability. Roughly 8% of tested cyber tasks were only solved at budgets of at least 10 million tokens, some at up to 50 million; a simulated cyber-range task estimated at about 20 human-expert hours was not solved by a tested agent until its allowance reached at least 30 million. This is direct evidence that hard evaluated tasks can have costly inference thresholds under particular models and scaffolds. A minimum observed successful budget is not a permanent floor for cheaper future models.
AISI’s fitted relationship between expert task duration and token requirements has substantial task variation and is supported on cyber and software-engineering evaluation suites, not internet-wide unauthorized operations. Newer models sometimes solve with fewer tokens, and larger budgets raise measured frontier task horizon sharply: the apparent trend itself depends on the evaluation cap. Numbers of tokens cannot simply be multiplied by today’s API list price into a detectable attack threshold; a successful serious campaign may find a lower-cost path, be guided by a human, or use an already discovered credential. Conversely cheap sandbox flags do not show that long-tail search, defended endpoints or persistence cost pennies. See the original compute-control hypothesis and the cross-pathway question. Do not post this technical budget argument as a second unread cyber item; keep for a triggered comparison.