AI review added candidate comments while missing most human-flagged changes
AI review added candidate comments while missing most human-flagged changes
Original: Crupi, Tufano and Bavota, Studying Quality Improvements Recommended via Manual and Automated Code Review, submitted February 12, 2026; accepted at ICPC 2026. They manually classified 739 change-request comments from human review of 240 PRs, then asked GPT-4 to review the same PRs. The AI suggested roughly 2.4 times as many changes as humans and found only about 10% of the issues humans mentioned. Roughly 40% of its extra comments concerned meaningful quality issues. This is complementary coverage, not a time or quality comparison of deployed review processes: human remarks are not a complete defect ground truth, and an extra ‘meaningful’ comment is not proof it improved the final software.
Read against DeputyDev's PR turnaround: faster first response can coexist with more suggestions to assess. An operator should record which warnings were accepted, rejected, fixed and checked after merge rather than treating AI-comment count or elapsed time as reviewer attention saved. Next compare measured active minutes and independently adjudicated issues for the same PR cohort. See review-clock synthesis.