Chen and Stratton: more agent code, longer review queue, no measured increase in issue closure

#topic

Chen and Stratton: more agent code, longer review queue, no measured increase in issue closure

Original: Fiona Chen and James Stratton, Artificial Intelligence in the Firm: Bottlenecks in Software Production. First version January 7; current paper August 4, 2026; data January 2021–March 2026; October 9 press discussion is not the date of the work. A collaborator at Jellyfish supplied anonymized engineering-event metadata for 718 opt-in client firms, about 300 million events (code, issues and meetings). Authors estimate firm-adoption differences against later and nonadopters, not a randomized agent-versus-human PR trial. Earliest API usage or agent bot/commit signature marks adoption; unsigned personal use and some integrations escape detection. They need early and late firms to follow comparable trends absent adoption; procurement timing may itself track management practices. They also check alternate event-study methods.

Measured stages. After agent adoption, per-worker monthly code metrics rose about 30% in lines, 20% in commits and 23% in PRs. Resolved issue and epic counts did not show statistically detectable gains; for issues, the 95% upper bound on gain was about 12% of baseline, not proof of zero output. Jira status is a team-reported completion proxy the authors interpret as deployed work, not verified user behavior, retention, value or a read of changed code. A text-derived issue-size check found no apparent shift, but it cannot guarantee identical task mix or desirability. No per-change token/model/CI bill was analyzed.

Review distinction worth preserving. The pooled agent-associated rise was 3.45 days against 7.03 days baseline, described as 49% longer ‘review time’; it is calendar time from PR submission to merge, not active reviewer minutes. Formal change requests rose about 12 percentage points against a 13% baseline, and comments rose 0.58 per PR against 1.66 baseline. The share of workers making at least one review rose four points from 29%, not a 14-point change nor a measured 14% increase in hands-on review hours. An imputation attributing elapsed time between signals (capped at two hours) to work categories and a separate 100-engineer Prolific survey are not task-level time sheets. The observed combination supports a downstream review constraint, but does not identify whether bad code, rising standards, increased arrival rate, review queues or extra iterations cause the delay.

AI review is still selective: by March 2026 nearly 80% of firms had used an AI review tool; 23.3% of comments were identified as AI-generated, and 10.8% of PRs received at least one AI-generated comment. The October 9 Ars account mistakenly treats the latter as the percentage of PRs authored by an AI agent. Agent adoption is firm-level evidence of any use, not that each PR has an agent author, and the review association is not a matched quality comparison. The paper's formal model treats coder throughput and the amount of work needed to verify drafts as complementary; it does not measure bugs, requester acceptance or causal review-minutes.

Disagreement and next test. A separate Zhou et al. company study (revised August 23) reports shorter peer-review time around rollout of an AI coding assistant for 200 developers; its abstract does not give an agent trial or the same outcome/design. It therefore challenges a blanket ‘AI always lengthens review’ story, not the distinct agent-versus-assistant association in Chen and Stratton. Pair firm-month throughput with one request-level attempts→review active time→independent acceptance→post-release follow-up ledger before treating a longer queue as a dollar or quality effect. Compare Microsoft PR output, ParallelPilot status, warning disposition and all-in costs without pooling endpoints.