From model bill to accepted-change cost: where is the missing denominator?
From model bill to accepted-change cost: where is the missing denominator?
Research brief and answer, October 10, 2026. If a router reduces agent spend per passing synthetic task, does it reduce the total expense and human attention for a maintained software change? A good answer follows one request across all failed attempts, code and tests, maintainer review, deploy, independent behavior check and subsequent repair, costing each step, with a fixed strong model and a human-led comparator. Not yet established. A new look at firm-level work finds a plausible downstream bottleneck but cannot turn calendar days into human minutes or per-change dollars.
Four subquestions searched: (1) does the routing preregistration count failed attempts, set-aside outages and the seeding cost? (2) is there a matched cost-per-independent-acceptance study? (3) is there an operator result joining actual human attention with spend and downstream output? (4) how should superseded, rejected and follow-up work count? Searched each once, then rechecked the October 9 press claim directly against its primary paper and the preregistration. A further primary 200-developer assistant study offers a different review-time direction, not a matched agent comparison.
Source list and findings. Kaiserauer's September experiment prices a synthetic task including every planned policy/trial cell; a session that hits a turn/budget cap counts as graded failure, but provider outages are set aside and rerun, with their spend recorded separately. The registered analysis excludes the cost of creating the parent seed as a sunk cost. This is a valid comparison of dispatch policies conditional on that starting point, not full cost from request. The paper's fixed Sonnet 5 worker was cheaper per pass than the router in an exploratory comparison, whereas the predeclared contrast was router versus always Opus 5. Qi's re-runs show that merged and tested performance fixes need not meet independent speed claims, but its selected 2025 PRs do not price reviewers or production impact.
Chen and Stratton follow 718 research-opt-in firms' events from 2021 to March 2026. Relative to staggered adopters, agent adoption was associated with 30% more lines and 23% more PRs, but no statistically detectable increase in resolved Jira issues/epics; for issues, their upper confidence bound was about +12%, not an exact zero. PR submit→merge time rose 3.45 days against a 7.03-day baseline, and formal change requests rose 12 percentage points. This is elapsed queue/iteration time, not active human review time or proof that agent-authored code is worse. They do not observe code contents; exposure is firm adoption, not assignment of each PR to a model. Their observed 10.8% is PRs receiving an AI review comment, not AI-authored PRs, despite the October 9 Ars retelling. Zhou et al.'s separate assistant-rollout abstract reports shorter peer-review time in one firm; interventions and endpoints differ, so no pooling or blanket claim that AI prolongs reviews.
Accounting protocol and unresolved gaps. Create one request ID and enumerate attempts (including rejected/discarded and rerun outages); cost model, caches, parent seed, test/CI, infra and external effects against that ID. Record review touch minutes distinct from calendar wait, approval/comment disposition, human integration and manual behavior checks. Link superseding PR and 30-day regression/repair to original ID; use a requester's confirmation and recontact window as separate outcome. For model-routing comparisons keep fixed strong baseline and same acceptance oracle; for firm economics do not turn extra submitted PRs or measured issue statuses into verified customer benefit. A notional cost per accepted maintained change is total attempts + human minutes valued at opportunity cost + CI/runtime + follow-up repairs divided by independently accepted, still-working changes; report request cohorts and unresolved ones rather than removing failures from the denominator. No source here joins these columns; stop speculative searching until a public same-request trace or a purpose-built trial surfaces. See paired-attention search and unit economics.