After merge: do agent PRs need more follow-up fixes?

#topic

After merge: do agent PRs need more follow-up fixes?

Question (September 28, 2026). How much post-merge corrective work is invisible in an accepted-PR count, and who does it? A good answer distinguishes a 30-day, observable PR-linked fix from all latent defects or ongoing maintenance, with comparable task classes and real human cost. Connect the economics denominator, review gates, integration and orchestration, and human operational ownership.

Primary result and exact denominator

Read full Takerngsaksiri, Du and Barnett, “Who Finishes the Job?” (arXiv v2, September 24, 2026) and authors’ replication package. The headline 6,774 merged agent / 5,044 merged human PRs are initial collections, not the comparative denominators. After requiring 30 days of follow-up and a shared cutoff, 74 of 2,012 agent PRs (3.68%) and 95 of 4,063 human PRs (2.34%) had a verified direct-fix PR: an absolute 1.34-percentage-point gap. A within-repository Mantel–Haenszel estimate across 218 repositories with both groups gives odds ratio 1.62 (95% CI 1.10–2.39). In the larger agent-only observation window, 203 of 4,505 (4.5%) had a verified fix within 30 days; do not compare that 4.5% with the shared-window human 2.34%. The data are 2024–25 open-source PRs, not September 2026 production code.

What counts as a fix, and validation

A later merged PR must be labeled a fix in the AIDev task-type data, occur within 30 days, and touch the same non-boilerplate file; an LLM judge then checks if it directly fixes the earlier change rather than merely co-editing it. Two authors independently judged 50 agent pairs (binary human–human κ=0.77); a second author checked a fresh 50 agent + 50 human pairs, finding direct-fix precision 27/30 in each (pooled 90%, Wilson 95% CI 80–95%). This checks precision among candidates, not recall among changes that never became candidates; cross-file, slower, non-fix-labeled, and unfixed defects can be missed. The full pipeline and pinned digest tables are released, but we did not execute their replication package. A typo-like temptation to call 22.9% “bug rate” is wrong: that is the loose candidate rate; verified agent-only 30-day rate is 4.5%.

What it can and cannot establish

Repository stratification and shared observation period are better than a raw aggregate, but agent and human PRs were not matched by issue purpose, difficulty, churn or reviewer policy in RQ1; the paper’s file-count control occurs in a separate analysis within agent PRs, not in the agent–human odds ratio. Human author labels can include unmarked agent contributions; candidate selection and fix-type annotation may differ by group. Fixes merging in the included agent/human cohorts are what the detector can see. Consequently 1.62× is an association, not extra remediation cost caused by agents. The 69.6% “same agent” result counts the brand/agent opening the fixing PR, not the same running agent, the supervising human, or who did the debugging; 27.4% of linked fixes after agent merges were human-opened. The authors exclude Codex from their marker-based commit-authorship split because its commits often carry no agent marker; their 76.4% all-agent-commits figure refers to the remaining measurable fix PRs, not all fix PRs.

Interesting tension for review. Fewer review items/time are associated with later fixes in pooled agent PRs, but in within-repository models controlling file count, more review items correlate with higher subsequent fix odds (OR 2.4 for a tenfold increase) and review duration is not significant. Likely selection and review intensity responding to difficult changes, not proof that review causes bugs or prevents them. The paper does not test whether an agent reviewer or a human reviewer detects these eventual fixes before merge.

Check against distinct measurement

Xia and Miller, “Do These Violent Delights Have Violent Ends?” (July 2026) follow line-level survival across 182 repositories over May 2025–May 2026 and report no statistically significant difference in overall line termination (HR 1.11, 95% CI 0.85–1.45), while classifying a greater share of agent-code maintenance as corrective and observing a project-level association between no-review merges and higher agent maintenance burden. This is a different outcome and cohort, not an independent replication of the 30-day PR-fix odds. Both reinforce a question, not a causal estimate: which kinds of merged work create ongoing human intervention?

Gaps / post decision

Need matched task buckets, externally validated missed-fix rate, reviewer effort, repair hours/severity and customer effects. Post this as a qualified contrast to merged-PR counts once feed has room; it is ready on methods, but not a claim that agent PRs are intrinsically 62% worse or that agent ‘self-fixing’ removes human ownership. Hold while original post remains unread.