ARTEMIS versus ten professional testers on a live university network: scoped exploitation, not botnet persistence
ARTEMIS versus ten professional testers on a live university network: scoped exploitation, not botnet persistence
Primary study: Lin et al., Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing, December 2025; revised 3 March 2026. University research team permitted ten security professionals and several AI-agent configurations to test the same 12 subnets (~8,000 hosts, seven public and five VPN) using issued student credentials and equivalent provisioned virtual machines. Researchers and university IT monitored activity, prohibited social engineering and destructive acts, and later remediated findings. Ten human testers were paid, recruited nonrandomly and allotted ten active work hours; two ARTEMIS runs lasted 16 hours but only their first ten hours were scored. This is a genuine human comparison on a real heterogeneous network, not a random sample of attackers or a test of unconstrained malicious behavior.
Results, correctly conditioned: the researchers' custom agent framework ARTEMIS used a supervisor, parallel sub-agents and vulnerability triage. Its multi-model configuration found nine validated vulnerabilities out of eleven submitted and placed second by the researchers' combined complexity/severity metric, ahead of nine of ten humans. A GPT-5-only configuration also submitted eleven findings, with six validated, and outscored five of ten humans; many standard agent scaffolds did worse and Claude Code refused the offensive task. All ten human testers found at least one critical vulnerability. ARTEMIS often stopped at a finding instead of pivoting deeper, reported spurious successes from redirected login pages, and missed a GUI-mediated flaw that eight of ten humans found. The claim that ‘AI beats expert hackers’ suppresses variance across configurations, one network, how the authors weighted complexity, and the higher false-submission rate.
Compute and defense: the cheaper GPT-5 run incurred $291.47 across 16 hours (~$18.21 in model API calls per hour); the second incurred $944.07 (~$59/hour). These are API costs, not a full criminal campaign including authorization, infrastructure, monitoring, identity, host persistence, and takedown; a researcher continuously watched each agent. University IT knew it was a test and approved flagged activity it might otherwise stop, so results do not measure evasion of a fully active defender. The cleanest discriminating replication would match human and agent time, access, defensive alerts and the same preapproved endpoints on multiple networks, then independently track verified exploitation, persistence, cost and response. Cyber pathway | compute/economics cross-area note. Feed: 29-page methods paper held for a weekend, conditional on cyber engagement or space opening; not another Tuesday long post.