Review queue, review touch time and accepted software: which clock did AI move?

#topic

Review queue, review touch time and accepted software: which clock did AI move?

Research brief and answer, October 11, 2026. Does agent-assisted coding increase reviewer effort, only increase queueing, or change how reviewers decide? A good answer distinguishes submit→first response, active reviewer attention, revision cycles, submit→merge and independent accepted behavior for the same request. Chen and Stratton report +3.45 calendar days submit→merge after agent adoption and more requests for changes, not active reviewer minutes or bugs. Across the comparative primary evidence below, no source supplies the paired review-touch→accepted-change ledger. Longer queue does not prove more reading; a shorter PR clock does not prove less reading; more gaze does not prove a better review.

Search path and sources. Searched once each for (1) the 200-developer AI-assistant counterpoint, (2) direct reviewer touch/quality measures, and (3) scheduling or reviewer-assistance interventions, then a gap search for active time and eye tracking. The Zhou et al. assistant study abstract reports shorter peer review, but the full text was not retrievable in this pass and the abstract does not expose the clock or assumptions; retain as a disagreement signal only. Read the first-party DeputyDev experiment, Khojah et al.'s controlled gaze study, Duma et al.'s GitHub review interactions and Crupi et al.'s matched review-comment comparison. Stop broad searching here: all four answer different columns and another aggregate count is unlikely to join them.

What each clock actually says. DeputyDev's July–August 2024 three-arm, filtered 721-PR field experiment measured reported review duration shorter with immediate AI feedback; authors give neither a clear active reviewer touch clock nor accepted defects and the tested version was one-pass GPT-4o, not the described multi-agent design. Khojah's August 2026 revision reports that 32 practitioners fixated longer on human-written code when falsely labelled AI-generated alongside a prompt, but scan length did not change detectably and deliberately injected faults were not scored. Label and added prompt confound which cue caused the attention. Duma's May 2026 descriptive GitHub study finds a very different kind of review: in the same repositories, observable human participation in agent- and human-authored PRs was 30.1% versus 30.8%, but agent PRs attracted fewer human-only reviews and more human instructions addressed to the agent. Silent reviews may be invisible; its ‘human review’ bucket includes ‘LGTM.’ Crupi's February 2026 paper compares GPT-4 with 739 human change-request comments on 240 PRs: more candidate AI comments, low overlap with human-flagged issues and some additional meaningful suggestions. It did not measure saved review minutes or downstream bugs.

Disagreement and inference. DeputyDev reports an improved elapsed PR interval while Chen and Stratton find a lengthened interval under firm agent adoption; different interventions, company/sample, dates (2024 test versus observations ending March 2026) and clock specifications prevent contradiction or pooling. Eye-tracked fixation is an active attention component but not full review effort, and the provenance cue—not real AI code—was assigned. GitHub comment absence is not absence of human reading. The shared operating lesson is measurement design, not a universal cost direction: for each incoming request, log PR arrival and first contact, reviewer foreground minutes (with an explicit off-screen/manual-check convention), actor and disposition for each comment, author response cycles, merge, independent scenario checked by someone other than the author, and 30-day repair. Randomize reviewer aid where safe or at least compare matched task difficulty and incoming PR volume, keeping acceptance and active minutes on the same cohort. The all-in accounting protocol should not price three days of waiting as three days of paid review.

Backlog: Khojah is a weekend long read if Dru reopens the reviewer-time question and the queue clears; Duma is a shorter weekend or weekday-with-standalone-summary counterpoint if ‘agents get no human review’ becomes a live claim. DeputyDev is a useful weekday review-clock counterexample after the queue clears, explicitly dated to 2024. None earns a fresh Sunday feed item while four already wait.