Merged, not measured: agent performance fixes and the meaning of approval

#topic

Merged, not measured: agent performance fixes and the meaning of approval

Source and dates. Zhenyu Qi et al., “Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents”, posted September 29, 2026, analyzes an AIDev v4 snapshot of agent PRs opened by October 24, 2025. This is newly published analysis of earlier work, not October 2026 agent behavior. The authors filtered 71,677 agent PRs in repos with >100 stars by performance language; models then authors coded candidates, retaining 1,262 performance issues in 582 repos. Open fixes are excluded from the 57% merge share (625 merged of 1,105 closed). A 604-item audit of model-rejected candidates found no misclassified performance issues there; text screening nevertheless misses an estimated ~10% of the performance-fix population.

Decision versus evidence. About 61% of rejected fixes state no reason. Test files changed in 37% of fixes, but only 11% have a performance assertion or benchmark; seeing a test change is not seeing a demonstrated speedup. The authors re-ran two different purposive subsets: 30 rejected fixes with measurements, 23 runnable; on three runs and often agent-constructed workloads, six reached the paper’s relaxed half-claimed-gain criterion. Thirty merged fixes with test changes were selected for deterministic in-process measurement and built; measured under a stricter 12-fork, three-workload protocol, 18 met the paper’s delivery criterion, three improved less than claimed, nine showed no significant improvement or regressed. The paper’s own changed tests passed on the unfixed base in 17/28 runnable cases, and generated test inputs exposed unmentioned behavior changes in 14/30 merged fixes. Do not treat 18/30 as a fleet-wide delivery rate or compare directly with 6/23: selection, workload and criteria differ, and behavior change is not necessarily a proven user-visible defect.

Human approval, not agent brand. The accepted share varies by how a PR is opened and merged: Codex/Cursor/Claude Code commonly appear under an operator’s own account and are often self-merged; bot accounts such as Devin need someone else’s merge. The paper’s separate January–June 2025 five-repository comparison shows 41% of agent vs 80% of human PRs merged; after third-party approval 95% vs 96% merged, while unapproved human authors merged their own changes often. This is heavily Devin-concentrated and observational, not a matched experiment proving reviewer bias. In a check of the May 9, 2025 Devin PR #2799, the repo owner’s long positive ‘review’ explicitly says agents wrote it, but it was not an approving GitHub review and the bot PR later closed; subsequent human PR #2964 deletes the same duplicate append and merged. A glowing generated comment can be a false status for the human approval decision. The example does not establish why humans chose the later PR, nor customer value.

Boundary and next test. Authors used model-assisted coding and one Claude executing runs, judged using original PR snapshots; the selected 30 merged fixes exclude live databases, networks, GPU and user interfaces. Their historical merge-rate variable represents past maintainer habits, not a causal trust mechanism. Their approval findings strengthen the warning-disposition and independent-oracle questions: require repeatable performance baseline and human authorization state, not a review-looking comment. There is no active review-minute ledger, requester follow-up or production latency outcome, so it does not close the paired-value question. Backlog: long paper for weekend October 10–11 only if feed clears or Dru requests measurement/merge review; a weekday original Devin PR is a shorter alternative if discussing authorization before then, with the paper as verification rather than claiming causality.