Unit economics of a maintained agentic change
Unit economics of a maintained agentic change
Origin and correction. On September 28, 2026, Dru asked to learn what people consider average/good unit economics per merged PR. After an initial answer emphasized lack of a universal $/PR and a pilot scorecard, Dru clarified: “I think you're being a bit nitpicky. Agreed with your points. cost per pr should have some bucketing by size/complexity, and we should also explore other relevant variables like bug rate, reversion, human time, code quality, etc. what i mean is i want to start researching and learning that area. how are people breaking this question down? common metrics/frameworks? what kind of results look good? etc.” Do not reduce this inquiry to designing Dru's internal pilot or repeat the benchmarking caveat as the answer.
Landscape for research
- Direct unit economics: spend per attempt and per merged/accepted change, token/inference and runtime distribution, human specification/prompt/review/rework time, success and abandonment, cost per task class/complexity; trace model/tool versions and repo context. Benchmark task costs are not production costs: Bai et al. (April 2026) find up to 30× token-usage variation between runs of the same SWE-bench Verified task and no monotone relationship between token use and correctness.
- Workflow and human capacity: eligible issues completed, merge rate, cycle time and review wait versus active effort, human intervention/edits, review burden, time diverted from other tasks. Microsoft observed +24% merged PRs associated with early CLI-agent adoption, but did not measure all-in costs/quality. GitHub's API has PR throughput/merge and consumption fields; its new review-stage series measures elapsed time for human-authored, human-reviewed PRs, not active human effort on bot-authored PRs.
- Post-merge quality: change failure, rollback/revert, post-merge fix incidence and severity, customer-reported bugs, test outcomes and code maintainability. DORA’s five delivery measures separate throughput (change lead time, deployment frequency, failed-deployment recovery) from instability (change fail rate, deployment rework rate). PR-level follow-up and maintainability need supplementary measures; inspect follow-up fixes paper before citing results.
- Developer experience and business value: SPACE/DevEx capture human flow and burden that throughput misses; product-side adoption, user satisfaction and successful feature outcomes test whether changes mattered. DORA recommends combining delivery and product or developer-experience frameworks, rather than changing measurement language wholesale for AI.
Types of 'good'. (a) More successful work in a matched size/complexity bucket with no higher reverts, defects or human effort; (b) less all-in cost or cycle time at equal quality; (c) a measured premium justified by stronger product outcomes. DORA’s 2025-research-based Quick Check offers industry comparisons for delivery throughput/stability, not an agent-specific $/PR target. DORA’s 2026 AI ROI model combines loaded salaries, time saved, AI/tooling and training costs, feature/revenue opportunity and change failure/downtime; inspect underlying assumptions and never report calculator example outputs as measured industry returns.
Research queue. Map common frameworks and competing vendor scorecards with definitions/denominators; collect actual benchmark ranges with task mix and quality scope stated, separate estimates from observed production outcomes, and look for studies that measure cost, human attention, acceptance and regressions together. Then give Dru concrete examples of strong versus weak results and the degree to which each generalizes.
Working accounting model, not a standard. Compare agent-assisted workflow with matched human-only work by task class and complexity. All-in cost per maintained accepted change = (agent and runtime spend on all attempts + active human scoping, prompting, review and correction + downstream rollback/fix costs over a stated window) / changes shipped and still healthy at that horizon. Report speed, quality and value separately; do not make this proposed composite the only lens.