Zhang et al. 2026: substantial digital biology uplift, not demonstrated physical uplift

#topic

Zhang et al. 2026: substantial digital biology uplift, not demonstrated physical uplift

Original preprint, v2 (13 March 2026), Scale AI/SecureBio/Oxford/Berkeley collaborators. An important update to the bio chain and to 2024 GPT-4 mild/non-significant information-only comparisons.

Design

57 biology novices: 10 non-STEM participants alternated internet-only and model conditions across written tasks; 47 STEM/Python participants were each assigned a condition for coding/agentic tasks. Several frontier mid-2025 models were available in the treatment, rather than one single model. Eight question/task suites cover virology, molecular biology and biosecurity-relevant problems, some lasting up to 13 hours. No real biological materials, lab operation, acquisition or dissemination. Some participants had model-evaluation experience; assignment scheme differs across cohorts.

What changed and what did not

The authors report 4.16× (95% CI 2.63–6.87) better overall benchmark accuracy; the detailed model describes this as a pooled odds ratio, with model-adjusted accuracy rising roughly 5% to >17%, not a measured 4.16-fold chance of making a weapon. AI users beat internet-only novices on seven of eight benchmarks and expert internet baselines on three of four available comparisons, although some differences in expert sampling and task conditions limit broad equivalence. Standalone models sometimes beat AI-assisted novices; easy access to a model does not imply optimal use. The 89.6% reporting little trouble accessing dual-use-relevant answers did not test bypass of every safety control or ability to acquire prohibited material. Some benchmark items appeared online; model choice changed mid-study; participants were not blinded. These are genuine evidence for knowledge and problem-solving uplift, not proof of wet-lab competence or a realized bioweapon. The original paper itself says in-silico uplift may not translate to physical-world effects.

Discriminator

Pre-registered tests with both internet and frontier-model arms, material acquisition tracked separately, and independent physical-lab validation by novices and already-skilled practitioners. Compare Active Site’s different endpoint, real-life actor reports and procurement controls.