METR’s 2026 repeat: an apparent reversal with a missing denominator

#topic

METR’s 2026 repeat: an apparent reversal with a missing denominator

Primary source. Joel Becker et al., METR, February 24, 2026, a report on the experiment begun August 2025, not a clean 2026 replication.

What changed. The February–June 2025 randomized issue-level experiment enrolled 16 maintainers doing 246 real tasks in mature repos with early-2025 tools and found AI allowed increased measured time 19% (95% CI +2% to +39%). New study: 57 contributors across 143 repos and over 800 tasks; 10 returnees, 47 newly recruited, $50/hr versus earlier $150/hr, with newer agentic tools. Among returnees, estimated AI-allowed task time was 18% shorter (CI 38% shorter to 9% longer); among new recruits, 4% shorter (CI 15% shorter to 9% longer). Both intervals include no effect; these are not causal population estimates for 2026 coding generally.

Why METR disavows its own new central estimate. Many developers declined to enrol because half their issues would forbid AI. In follow-up surveys, 30–50% said they withheld at least some tasks because they did not want to risk doing them without AI. Rate changes and incomplete/disproportionate work on no-AI tasks further distort representativeness. Parallel agent use makes manual time-spent estimates ambiguous. Task mix, output quality and the relationship of saved time to value can change with availability of agents. METR thinks newer tools likely help more than the early-2025 tools, but says these data weakly constrain by how much. Neither selectively submitted tasks nor the sample automatically imply a calculable lower bound without assumptions.

Design implication. For an engineering organization, record eligible tasks and rejected/deferred tasks before randomization, as well as quality and review time, or a randomization within a willing subset may describe precisely the tasks where it is least painful to turn the agent off. METR proposes higher-paid shorter trials, fixed tasks, observation, questionnaires, or developer-level randomization (which can worsen developer selection).

See Microsoft’s 2026 agent rollout for a larger observational result with a different endpoint, and the evidence question.