Physical-laboratory uplift tests: the endpoint defeats an easy headline

#topic

Physical-laboratory uplift tests: the endpoint defeats an easy headline

Original studies: Hong et al., pre-registered 153-participant randomized study, fielded June–August 2025, published February 2026; Romero-Severson et al., Los Alamos 10-person pilot (2025, updated 2026); a contrasting December 2025 benign experiment with GPT-5 and trained scientists. See mechanism synthesis.

Large pre-registered trial

77 novices assigned access to several mid-2025 frontier models plus internet, 76 internet only, in supervised BSL-2 lab tasks modelling segments of a reverse-genetics workflow. Pre-registered core task-sequence completion: 4/77 (5.2%) AI versus 5/76 (6.6%) internet; risk ratio 0.79, 95% CI 0.24–2.62, p=.759. This is not evidence AI lowers success, and does not rule out moderate uplift. AI users progressed numerically more in most intermediate steps; a post-hoc model estimates a 1.42-fold increase for a hypothetical average lab task, credible interval 0.74–2.62. Researchers expected higher success when planning and note study was underpowered after unexpectedly low completion. No material acquisition/infrastructure setup; individual tasks decoupled and some simplified; no end-to-end threat demonstrated or ruled out. Presenting this as ‘AI can’t help in the lab’ would misstate the data.

Small pilot and productive expert pathway

Los Alamos randomized five novice employees per arm to internet versus internet plus o1 in a supplied, safe lab exercise. On two allowed attempts 4/5 AI and 3/5 internet completed; exploratory pilot too small for inference. Human experts offered help at impasses in both arms; all equipment and materials were supplied. The investigators observed model-interface mistakes and long responses slowing progress. In a different population and endpoint, OpenAI/Red Queen Bio’s controlled, benign cloning study used GPT-5’s proposals, human scientists performing lab work, and experimental feedback across rounds to improve a protocol 79-fold over the chosen baseline (three validation runs). This demonstrates useful AI-assisted wet-lab innovation for experts on a specified assay, not novice biothreat production or universal 79-fold research speedup.

Interpretation

Large digital-knowledge uplift and null primary physical-endpoint results measure different links; the former must not be projected to material success, and the latter must not become a claim of an enduring skill barrier as models, interfaces and operators change. Next: original measured expert-uplift and exposure to procurement controls, not another biological QA benchmark.