Agentic code review: from human gate to evidence and bounded autonomy

#topic

Agentic code review: from human gate to evidence and bounded autonomy

Why here. On September 28, 2026, Dru added a specific research interest: emerging code-review best practices, what moves quality and effort, who actually uses zero human review, in what contexts, and how confidence is built. Investigate, do not presume a universal zero-review rule. Cross-link review labor/economics and post-merge quality.

A real autonomy case, with limits

OpenAI, Harness engineering (February 11, 2026): an internal-beta/externally-alpha product started from a new repository in August 2025. Three engineers drove ~1,500 merged agent-written PRs over five months; the post says human PR review is optional and increasingly handled agent-to-agent; agents often merge their own PRs. Not zero human judgment: humans specify priorities/acceptance criteria, inspect outcomes, feed user reports into tooling, and intervene for escalations. Confidence comes from legible repo-local documentation, strict architectural dependency boundaries, custom linters and structural tests, agent review loops, worktree-local runnable UI, screenshots and app-driving, traces/metrics, and cleanup agents. The authors acknowledge app breakage, QA as a bottleneck and special repository investment, explicitly warning against broad generalization. Its time-savings estimate is author-reported, not a matched quality outcome. Distinguish OpenAI’s separate Symphony workflow, whose authors explicitly say humans review agent results, from the optional-code-review case.

Specification-first convergence, August 2026 is a second claimed zero-human-code-review case on a large TypeScript architectural refactor, but a single report with no pre-existing test oracle; inspect full protocol, independent verification and follow-up before using as proof of general reliability.

What has and has not moved the needle

GitHub engineering, July 2026: merely giving Copilot code review better general search tools worsened its internal quality/cost benchmark; tuning it to start with the diff, gather targeted surrounding evidence and not browse broadly reduced average review cost ~20% while maintaining measured review quality. Vendor internal benchmark, not proof of fewer incidents or reduced human review time. GitHub September 11 ensemble update reports +47% high-severity comments addressed per review and ~8% lower review cost in its experiments; addressed comments are a proxy, not ground-truth bugs averted. c-CRAB independent code-review benchmark, March/April 2026: agents collectively solve only ~40% of its tasks derived from human reviews (not a direct production bug-miss rate); suggests complementarity and need for independent validation.

GitHub September 1, 2026: Copilot code review may submit a merge-counting approval if admins enable it, configurable by repository and file path; the default is off; approval assessments alone do not satisfy merge rules. Investigate the role of path-specific risk policies, independent reviewer agents, deterministic tests/static analysis, release controls, canaries and revert paths in any zero-human approach.

Critical distinction. Code review of diffs is not the same as approval of tool actions: OpenAI’s Auto-review (April 2026) applies an independent agent to sandbox-boundary tool calls, cutting synchronous human tool-approval stops; do not present its ~200× reduction as a PR-review result.

Research agenda. Compare practices on defect recall/precision, severity of confirmed findings, addressed rate versus correctness, human minutes/interruptions, PR lead time, reversions/incidents, and longitudinal maintainability. Find matched or prospective production measurements rather than inferring safety from green CI or merge speed. Sort autonomy by change risk (docs, isolated fixes, core invariants, authentication and data migration), blast radius, recoverability and ownership. Seek examples and failures from organizations besides OpenAI, GitHub and vendors.