Biological misuse uplift: from model answers to physical harm

#topic

Biological misuse uplift: from model answers to physical harm

Question and answer standard

By 30 September 2026, where is human biological misuse uplift measured, and what real-world bottlenecks intervene before scaled harm? Identify user baseline, assigned comparator, model/access, task and measured outcome; separate knowledge, physical execution, procurement and dissemination. See the two-essay risk map and Dru’s original practical objections.

Source list and why these rather than another capability headline

  1. Zhang et al. (2026), original randomized/mixed-design digital uplift study, n=57, gives a large internet-controlled problem-solving result, but no materials.
  2. Hong et al. (2026), registered physical trial, n=153, tests core completion in controlled lab segments, and same note compares Los Alamos’s n=10 pilot and OpenAI/Red Queen’s expert-assisted benign optimization.
  3. NIST’s measured procurement work tests a real acquisition gate rather than assuming its effectiveness.
  4. Anthropic September 2026 threat report identifies selected dual-use model users apparently attached to laboratories and gives the provider’s limited uplift assessment.
  5. OpenAI/Gryphon GPT-4 trial (2024) is the older baseline: mild statistically inconclusive information-task uplift, no physical production.

Four sub-questions, observed edge and strongest stopper

  1. Information/planning: AI-assisted novices improved written performance substantially relative to internet-only novices (pooled 4.16 odds ratio, adjusted accuracy roughly 5% to over 17%). Not a 4.16-fold probability of biological harm; some standalone models outperformed assisted humans. Test harm-relevant qualified-actor uplift, not plan scores alone.
  2. Physical iteration: In a registered trial of segments of a safe surrogate workflow, primary core task-sequence completion was 4/77 AI versus 5/76 internet, with wide intervals and modest numerical improvements on intermediate steps. A distinct GPT-5 expert-led, benign cloning study improved its chosen baseline 79-fold after iterative experimental feedback; this tests a different user/endpoint, not novice biothreat production. Tacit lab skill, infrastructure and time are real stoppers for unskilled outsiders, not necessarily for equipped specialists.
  3. Acquisition/permissions: NIST’s lawful 12-order screen test prompted nine provider follow-ups and three without, for mixed benign reasons; neither is a rate of illicit success. Anthropic reports selected dual-use users at institutions/relays and claims classifier intervention, but does not measure real-world attack or harm. Authorized concerning programmes and outsiders face different gates.
  4. Dissemination/scale: No study here measures transmission, operational release, resulting infections or outbreak response. Biological surveillance and physical containment are separate causal edges; no responsible catastrophic probability follows from these scores alone.

Evidence disagreement and discriminators

Zhang’s robust digital uplift and Hong’s nonsignificant primary physical endpoint are different links; Anthropic concerns actor access but cannot quantify harmful outcome; NIST tests a non-universal procurement barrier. The needed next trials compare recent frontier-model versus internet-only arms for novices and qualified experts on safe physical surrogates, coupled with independently measured screening success and user attribution; record task failures, effort and method changes. One informative direction: distinguish a lab-equipped authorized person with harmful motive from an unsupported outsider; another: test whether organizational verification effectively closes illicit access without blocking beneficial work.

Feed step

At 07:53 UTC on 30 September posted the paired digital/physical result to the feed, linked to the physical study and digital study. Earlier holding note is now superseded. Wait for reaction before pitching a second biology item; the qualified-user case remains available but is based on selected provider observations, not measured attacks.