What do productivity studies of coding agents actually measure?

#topic

What do productivity studies of coding agents actually measure?

Question. Are agents making engineers faster on consequential work, and for whom? A good answer separates randomized developer/task experiments, observational production throughput, surveys and benchmarks; names the tool era, denominator and unit of output.

Research brief completed September 28, 2026. (1) Primary randomized field experiments on experienced engineers, including a repeat: METR repeat. (2) Strong controlled gains with different worker/tool population: Cui et al. workplace Copilot trials. (3) Real deployment telemetry and limits: Microsoft CLI-agent rollout. (4) Temporal transport: 2023 autocomplete, 2025 early agents and 2026 CLI agents are not one treatment; the '2026' publication date on Cui et al. is not a 2026 experiment.

Short primary-source list. METR experimental redesign; METR 2025 original RCT; Cui et al. published RCTs and full author draft; Murphy-Hill et al. 2026 enterprise agent field observation; METR 2026 technical-worker survey; METR human vs automated scoring.

What the evidence now says. METR early-2025 issue-level RCT finds 19% longer work times in 16 experienced maintainers; its late-2025 repeat offers point estimates suggesting improvement but METR says the task/developer selection makes magnitude unreliable. Cui et al. find +26.08% PR-based completed tasks in randomized access to 2022–23 code completions, especially newer engineers. Microsoft authors find +24.0% merged PR throughput associated with adoption of 2026 CLI agents; observational early adopters and PR-count endpoints cannot establish equal gains in shipped value. These studies differ in treatment, worker, timeframe and measured outcome, rather than simply contradicting each other.

Still missing. An agentic-era randomized production study that captures all eligible work, agent spend, review time, defects and customer outcome. Also longitudinal knowledge/maintenance costs: Anthropic’s small learning experiment finds lower immediate mastery in a junior/new-library task; cannot directly establish long-term harm. A September 22, 2026 paper on follow-up fixes in merged agent PRs is a candidate quality countercheck; read full methods before treating the odds ratio as causal.

Next question. Define 'done' as mergeable change and subsequent maintenance rather than passing a test: METR’s 2025 18-issue exercise found 38% passing maintainer tests but none of 15 manually reviewed PRs mergeable as-is (early model, selected tests, small sample).