DeputyDev's shorter PR turnaround is not measured reviewer minutes

#topic

DeputyDev's shorter PR turnaround is not measured reviewer minutes

Original: Khare et al., DeputyDev — AI Powered Developer Assistant, posted August 13, 2025; TATA 1mg trial ran July 27–August 27, 2024. Their PR-review assistant brings together Bitbucket diff, Jira request and Confluence approach, searches related code, and provides immediate AI feedback. The full described multi-agent design uses separate quality concerns and a consolidation step, but the reported experiment used a comprehensive one-pass workflow with GPT-4o—not that multi-agent implementation. Do not infer a measured gain for the agents.

The authors say they allocated roughly a third of PRs per repository to each of two control groups and one treatment, filtering out the smallest 10% and largest 25% by changed lines and repositories without enough PRs in all three groups. After filtering, 244, 238 and 239 PRs remained. Mean ‘review time’ was 239.57 and 278.14 hours in controls and 197.97 hours in treatment; medians were 0.76 and 0.78 versus 0.41 hours. These huge mean/median gaps and the proposed mechanism of immediate automated first feedback point to a workflow/elapsed interval, not minutes a human read a diff. Their methods do not give a clock definition precise enough to separate first response from merge and do not give individual reviewer touch-time traces. The paper claims significance, but reports no independent reviewer-accuracy or customer/release outcome table and gives no cost per accepted change. Allocation is described, not clearly documented as randomized.

This is a genuine counterexample to ‘an AI reviewer necessarily slows the PR clock,’ under a different intervention and a 2024 task mix; it does not negate the later firm agent-adoption review-queue association. Nor does it prove that review labor fell: author waiting, active reviewer attention and discovered faults must be tracked separately in the review-clock brief.